LONG CONTEXT NEEDS A SESSION BUDGET, NOT JUST ENOUGH RAM FOR THE WEIGHTS
WHAT THE RESEARCH ACTUALLY SHOWED
REPORTED: The 2023 PagedAttention paper reported 2–4× throughput against FasterTransformer and Orca at comparable latency in its experiments. We have not reproduced those results; they are directional, not a forecast for another runtime or Mac. Its block allocation addresses fragmented memory and duplicated KV state. Primary paper, method and evaluation.
THE MEMORY BUDGET GROWS WITH THE WORKLOAD
ESTABLISHED: weights persist across requests, while cached attention keys and values depend on the tokens retained for each active sequence. More simultaneous sequences therefore need additional state; longer outputs can keep expanding it after the prompt finishes. Paged allocation reduces waste and permits sharing eligible blocks, but does not eliminate the state associated with distinct tokens. PagedAttention: memory challenges.
ESTABLISHED: vLLM's current tuning guide describes request preemption when KV space runs out, followed by recomputation when capacity returns. It identifies concurrency and batched-token limits as controls. The operational inference is to budget the longest useful prompt, intended output and concurrent active work together, then examine queueing and preemption rather than treating successful model loading as acceptance. vLLM tuning guide.
UNIFIED MEMORY CHANGES ACCESS, NOT THE NEED FOR HEADROOM
ESTABLISHED: MLX allows CPU and GPU operations to access arrays in the same memory pool without explicit device-to-device copies. That describes where tensors live; it does not establish how many long sessions fit. Our capacity-planning inference is to leave room for cached state and temporary work instead of subtracting the weight file from the machine's advertised RAM and calling the remainder guaranteed context capacity. MLX unified-memory documentation.
ESTABLISHED: MLX LM documents a bounded rotating cache, reusable prompt caches and configurable prefill steps. These solve different problems: smaller prefill steps reduce peak prompt-processing memory at a speed cost; a smaller rotating cache trades retained context for memory and can reduce quality. Reusing a cached prefix avoids repeated prompt computation, not the cost of every unique continuation. Check the installed command's help before choosing options. MLX LM documentation.
READ THE RUNNING CONFIGURATION BEFORE INCREASING CONCURRENCY
ESTABLISHED: llama.cpp documents context size, parallel slots and unified KV settings separately. Its read-only properties endpoint exposes the build, slot count and generation settings. On your own already running single-model server, inspect these values; do not assume a context number means that allowance for every simultaneous request. The commands below neither submit generation nor change configuration. llama.cpp server reference.
llama-server --version
llama-server --help
curl --fail --silent --show-error --max-time 10 \
'http://127.0.0.1:8080/props?autoload=false' \
| jq '{build_info, total_slots,
n_ctx: .default_generation_settings.n_ctx}'
ESTABLISHED: use the configured local port and authentication where required; an unavailable field is unknown, not zero. For MLX LM, inspect mlx_lm.generate --help and the existing launch configuration. For vLLM, review existing preemption metrics or logs before changing limits. These observations identify configuration and pressure; they do not measure latency under your workload. MLX LM help · vLLM observations.
WHERE PROXIES.SX FITS
OURS: the public tiers API, checked September 28, 2026 UTC, describes Starter as 27B-class models at 4-bit (16K context)
. The current catalog API publishes pinned revisions, quant, requiredMemGB and maxContext; its three entries specify context limits of 16,384 or 32,768. These are catalog settings, not measured concurrency or throughput. Fetch them again when planning:
curl --fail --silent --show-error --max-time 10 \
https://api.proxies.sx/v1/peer/compute/catalog \
| jq '.models[] | {id, revision, quant, requiredMemGB, maxContext}'
curl --fail --silent --show-error --max-time 10 \
https://api.proxies.sx/v1/peer/compute/tiers
OURS: review the compute marketplace and current service limits before selecting a node. Our documented service permits one running job per node; a general-purpose server's parallel-slot feature is not a promise about this rental API. This article certifies no available stock, benchmark, session count or workload performance.