The useful limit is accepted work within a budget, not the largest checkpoint that starts.
The biggest model a machine can load is a poor default for the work that machine needs to finish. September kept returning to the same distinction: a checkpoint can fit, emit tokens, and still lose the working session. The useful model is the one that delivers an accepted result before its memory traffic, prompt processing, or recovery cost breaks the workflow.
Five lessons survived the month’s tour from text inference to local video. Budget the active working set. Treat sparse models as a placement problem. Choose a runtime for the actual hardware. Reserve memory for context and cache. Apply the same discipline to visual graphs. Recent releases sharpened those lessons without producing a universal winner, and their dates matter because the implementation beneath a saved configuration keeps moving.
The working set became especially concrete in Hugging Face’s September 22 announcement that Transformers can run packed GGUF weights through compatible ggml-derived Metal kernels. Its Qwen3.5-4B example lists an 8.42 GB BF16 file and a 2.74 GB Q4_K_M file. Those are storage sizes, not peak-memory measurements. Without a compatible quantization kernel, the loader falls back to dequantization and uses more memory (Hugging Face announcement). The runtime decides whether the smaller representation remains smaller during inference.
That path initially targets Apple Silicon and the Qwen3.5 architecture, requires compatible PyTorch and kernel builds, and the announcement specifies Transformers main pending a release. A Python import succeeding does not establish that the intended packed kernel ran. Verify the backend and inspect allocation during the complete task before crediting the file-size saving. The useful question is which tensors occupy fast memory while the next operation executes.
Sparse models make that question harder. Qwen’s official Qwen3.6-35B-A3B card specifies 35 billion total parameters and 3 billion activated (Qwen model card). Activation limits the work selected for a token; it does not delete the remaining weights. The full expert collection still needs storage, and the selected experts need a route to execution. A long prompt can exercise a different mix of experts from a short continuation, so a good decode result leaves the first large prefill unresolved.
Colibri’s September 24 v1.12.1 release offers a more useful placement lesson than a giant-model loading screenshot. Its automatic dense-trunk placement now times a matrix-vector operation on CPU and GPU before committing. The reported Tesla M10 configuration ran the placed components slower on GPU, so automatic placement withdrew and returned that VRAM to the expert cache. Explicit placement remains available (Colibri v1.12.1). That is project-reported evidence on a particular machine, not a speed result reproduced here.
The same release repairs Qwen3.6 punctuation and whitespace tokenization against the reference and restores default-context requests by clamping the output ceiling to the room left after the prompt. Those fixes belong beside the placement change. An engine that finds capacity while tokenizing code differently or rejecting an ordinary request has not delivered the model’s intended service. I would rerun representative code prompts and context-limit fixtures before promoting the update.
The runtime lesson reached Apple Silicon through MLX v0.32.3, released September 29 UTC. Its notes include M5 Ultra tuning for non-quantized matrix multiplication, a sorted quantized-gather overflow fix above 32K, and changes to attention kernels and GGUF tensor-dimension validation (MLX v0.32.3). These are concrete changes in the execution layer. They do not establish a throughput multiplier for another Mac, another quant, or an MLX-LM environment that has not been tested with that core version.
Runtime choice is therefore a hardware decision with a maintenance bill. A convenient serving surface, a compatible model format, and a fast kernel path all matter, but none substitutes for the others. I would start with the simplest supported stack that meets the task, then move to a specialized engine only when a measured shortcoming earns the added dependencies. Rebuilding an environment to chase a chart is expensive when the chart omits the prompt shape that dominates your day.
The portable baseline moved too. llama.cpp v0.5.0, released September 23, includes backend and parser fixes, enables CUDA graphs for multi-token-prediction drafting, and repairs server router lifecycle behavior (llama.cpp v0.5.0). A saved command from early September can still run while exercising a different implementation. Preserve the executable or commit alongside the model hash and launch settings. “Same model” is too little information to explain a changed result.
Context consumes the memory left after those choices. The v0.5.0 server documentation exposes separate key and value cache dtypes, a host prompt-cache budget, context checkpoints, and shared or separate sequence buffering (tagged server documentation). They are different controls over different state. A smaller weight quant does not automatically make a long context affordable, and retained prompt snapshots consume memory even when they save future computation.
Consider a hypothetical coding workflow that accepts a model at short context, then adds a repository excerpt and a second concurrent request. The acceptance configuration changed twice. Record the deepest intended prompt, the output allowance, and the concurrency ceiling together. Test the first uncached turn, a continuation with a stable prefix, and a tool result that changes the context. Cache reuse can hide an expensive prefill until the edit that matters invalidates it. Hybrid and recurrent architectures also need their own accounting; one conventional attention-cache formula cannot price every model.
Visual graphs follow the same hierarchy. ComfyUI v0.37.0, released September 21, added automatic fast-disk detection with Aimdo 0.5.5, a switch to disable that path for debugging, and a lower Wan peak-VRAM path for comfy-kitchen attention (ComfyUI v0.37.0). Those release items concern residency and execution. They do not prove that a particular image or video graph fits, runs faster, or preserves quality on your workstation.
The graph gives you useful boundaries: text encoding, repeated denoising, latent decode, and finishing. Diffusers documents the cost distinction between moving whole components and sequentially offloading their smaller submodules. Sequential offload can save more memory while adding repeated transfers and substantial latency (Diffusers memory guide). Moving a component once at a stage boundary deserves a different test from moving blocks during every sampling step. Larger canvases and longer clips still grow the intermediate work after weights have been compressed.
A second card is another place to assign work. Native ComfyUI’s tagged source includes device-selection nodes for models, text encoders, and VAEs (device-selection source). Start with a whole-component boundary or two workers joined by verified files. Require an improvement in completed-job time before accepting a split that moves frequently used tensors across devices. More occupied VRAM is an observation; fewer failed or delayed jobs is the outcome.
The strongest objection to all this budgeting is that a larger, slower model may produce an answer the smaller one cannot. That is a valid reason to accept latency. Overnight extraction and occasional difficult analysis can justify a configuration that would ruin interactive coding. Write down that choice as a task budget, including retries and the operator’s review time. A slow model earns its place through accepted work that the faster alternative misses.
I would end a model-selection experiment with three separate verdicts. Fit records peak GPU and host memory at the intended workload. Speed records cold start, prompt processing, cached continuation, and completed-task time. Quality uses the same representative acceptance set before and after a precision, placement, or runtime change. Keep failures and raw settings beside successes. None of the release notes above is a new local benchmark, and September’s configurations deserve no permanent exemption from those checks.
Pick the strongest model that meets the workflow’s latency and quality floor. A machine that prints one token has proved it can begin. Choose the configuration that can finish.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


