Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions QUICKSTART.md
Original file line number Diff line number Diff line change
Expand Up @@ -223,6 +223,8 @@ ollama show qwen2.5-coder:7b --modelfile # the FROM line shows ...blobs/sha2
quantprobe bench --gguf ~/.ollama/models/blobs/sha256-<hash> --total 7.2 --active 7.2 --bits 4.5 --vram 0 --ram 32 --ram-bw 45 --disk-bw 2
```

**`quantprobe audit-ollama --measure` reads the placement of the model you named, exactly.** It takes the `ollama ps` row whose NAME *equals* that model — tag included — and an untagged name resolves to `:latest` the way `ollama run` resolves it, nothing more (in `hub.example:5000/team/model` the colon is the registry's port, not a tag; the tag, if any, is in the last `/` component). `qwen2.5:7b` never answers for `qwen2.5:14b` or for `qwen2.5-coder:7b` — different weights, different layer splits — and `:latest` is a default, not a wildcard over your other tags. When no row matches, because the model is not loaded or ollama prints a name in a form this does not recognise, **the placement stays unknown and no comparison is printed**: a neighbour's split benched at a neighbour's context depth would be a confident `Nx AVAILABLE` number about a model you did not ask about, which is worse than no number.

(Windows: `C:\Users\<you>\.ollama\models\blobs\`.) Two Ollama gotchas the law keeps exposing in the wild: Ollama's **default context window (~4k) is smaller than coding-agent payloads** — Continue/Cline overflow it, truncation breaks prompt-cache reuse, and every request re-prefills from zero (minutes on CPU). Either raise it (`PARAMETER num_ctx 16384` in a Modelfile / `OLLAMA_CONTEXT_LENGTH`) or serve with `quantprobe run --serve --extra "-c 16384"` instead. And on CPU-only boxes, prefer MoE models — dense 14B decodes ~3 tok/s where a 30B-A3B does ~8 with more intelligence (`plan` shows this per-box).

## Recipes — getting the most out of it
Expand Down
30 changes: 29 additions & 1 deletion quantprobe/ollama.py
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,29 @@ def ollama_bin():
return shutil.which("ollama") or shutil.which("ollama.exe")


def _is_same_model(row_name, want):
"""True when an `ollama ps` NAME cell addresses the SAME model as `want`.

Whole-name equality, because every looser rule has a counter-example on a real store:
a tag is part of the identity (qwen2.5:7b and qwen2.5:14b are different weights with
different splits), a name is not a prefix namespace (qwen2.5-coder:7b is not qwen2.5),
and a registry host may carry its own colon (registry:5000/ml/mistral:v3). The one
normalization is ollama's own default tag: `ollama run qwen2.5` loads qwen2.5:latest,
and that is what ps prints.

Whether `want` is already tagged is read off the LAST slash-separated component only,
because that is the only place a tag can appear: in hub.example:5000/team/model the
colon belongs to the host's port, so treating any colon as a tag leaves that name
permanently untagged-but-unnormalized and its resident :latest row unreachable.
Nothing else about the namespace is rewritten - no host is added, stripped or aliased -
so two names that differ anywhere before the tag stay different models.
"""
if row_name == want:
return True
tagged = ":" in want.rsplit("/", 1)[-1]
return not tagged and row_name == f"{want}:latest"


def loaded_placement(name):
"""What ollama ACTUALLY chose: (gpu_percent, ctx) from `ollama ps`, or (None, None).

Expand All @@ -81,6 +104,10 @@ def loaded_placement(name):
placements and comparing them would look like a speed claim while actually being a
category error. Measured here on a 6 GB card: ollama loaded qwen2.5:7b as 16%/84% CPU/GPU
even though the model fits, so the gap was never evidence about anyone's prediction.

The row has to be THIS model's. --measure feeds gpu_percent into llama-bench's -ngl and
ctx into -d, so a neighbouring row benches one model's layer split at another model's
depth and prints the difference as a placement recommendation.
"""
b = ollama_bin()
if not b:
Expand All @@ -94,7 +121,8 @@ def loaded_placement(name):
import re

for line in out.splitlines():
if not line.startswith(name.split(":")[0]):
cells = line.split() # NAME is the first column; the header and blanks fall out here
if not cells or not _is_same_model(cells[0], name):
continue
m = re.search(r"(\d+)%/(\d+)%\s*CPU/GPU", line)
gpu = int(m.group(2)) if m else (100 if "100% GPU" in line else None)
Expand Down
36 changes: 36 additions & 0 deletions tests/smoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -2846,6 +2846,42 @@ def t_ollama_eval_rate_is_generation_not_prompt():
return None


def t_ollama_placement_reads_the_row_for_the_model_it_was_asked_about():
"""audit-ollama must read ITS model's `ollama ps` row, not a neighbour's.

The match was `line.startswith(name.split(":")[0])`: it threw the tag away and then
prefix-matched, so with qwen2.5:14b and qwen2.5:7b both resident, asking about the 7b
returned the 14b's split and context - and with qwen2.5-coder:7b listed first, a model
that is not even the same weights. That number is not cosmetic: --measure feeds gpu_pct
into `-ngl` and octx into `-d`, so the wrong row benches one model's layer split at
another's depth and prints the result as a placement recommendation.

The cases live in tests/test_ollama_placement.py - the tags, the prefixed sibling, the
registry host that carries a port, the three PROCESSOR shapes, the pre-CONTEXT layout,
the not-loaded and no-daemon paths - and this hook runs ALL of them through that file's
run_smoke(), which needs no pytest. Copying a few assertions down here instead would be
a second copy of the intent: every case added there afterwards would be invisible to
`python tests/smoke.py`, which is the same way a test once sat below this runner and
never executed. A non-empty return from run_smoke() is a failure, every one of them.

Fixtures are `ollama ps` output shapes with the subprocess boundary stubbed; no daemon,
no model, no network."""
import importlib.util

path = os.path.join(os.path.dirname(os.path.abspath(__file__)), "test_ollama_placement.py")
spec = importlib.util.spec_from_file_location("qp_ollama_placement_cases", path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)

cases = mod.smoke_cases()
assert len(cases) >= 14, (
f"only {len(cases)} placement case(s) collected from {os.path.basename(path)} - the "
f"suite shrank, which reads as green while covering less")
failures = mod.run_smoke()
assert not failures, f"{len(failures)}/{len(cases)} failed: " + " | ".join(failures)
return None


def t_ollama_store_reader_survives_a_broken_store():
"""audit-ollama reads a directory it does not own, so it must degrade rather than crash.

Expand Down
Loading