Benchmarks
Measurements, with the configuration attached.
A throughput number without the hardware, the build and the load it was taken under is not a measurement. Everything below is reported with the conditions that produced it.
Single request
What one un-contended request decodes at, and prompt processing. This is the figure a client experiences before any other load is present.
| Model | Single-request decode | Prompt processing | Benchmark context | Grade |
|---|---|---|---|---|
| Qwen3.8-27B | 63.70 tok/s | 1130 tok/s | 4,096 ctx | MEASURED |
Decode vs concurrency
All streams fired simultaneously, 256 completion tokens each. Aggregate is total output across every active stream; per stream is what one of them experiences once the GPU is shared. Measured at 4,096 context with an f16 KV cache and short prompts — concurrency with long live contexts has not been measured, and a long context both takes KV memory and slows each decode step.
| Model | Streams | Aggregate | Per stream | VRAM | Serving cost / M out |
|---|---|---|---|---|---|
| Qwen3.8-27B | 1 | 63.68 tok/s | 63.70 tok/s | 15,298 MiB | $0.779 |
| 2 | 81.19 tok/s | 40.83 tok/s | 15,892 MiB | $0.611 | |
| 3 | 114.22 tok/s | 39.03 tok/s | 16,528 MiB | $0.434 | |
| 4 | 127.19 tok/s | 32.47 tok/s | 17,102 MiB | $0.390 | |
| 8 | 167.20 tok/s | 21.73 tok/s | 19,550 MiB | $0.297 |
Aggregate does not scale linearly with concurrency: measured throughput reached 2.6× at eight streams, not 8×. Per stream therefore falls as concurrency rises, and the two columns answer different questions — one is capacity, the other is the latency a user feels. Serving cost is derived from the measured host rate and aggregate throughput; it is the cost basis a customer price is set from, not a published price.
Serving configuration
The build and the flags are ours and are recorded, because a host choosing them would make every number above unreproducible.
- GPU
- NVIDIA GeForce RTX 3090 (24,576 MiB)
- Host
- Vast.ai (verified) · US / SE (two runs) · $0.1786/hr
- Runtime
- llama.cpp 0.5.0-dev (build 11235, commit 6c7a87f7e)
- Driver
- 595.71.05, CUDA 13.2
- Batch / ubatch
- 2048 / 512
Artifact
The exact weights served, identified down to the revision hash.
- Quant
- UD-IQ4_XS
- Size
- 14.25 GB
- Upstream
- unsloth/Qwen3.8-27B-GGUF
- Revision (sha256)
- 40fac4050e940397dbf13087afd50f4734a11805bf9d65ef8ddd7483470e6199
- License
- apache-2.0
Context and KV
The maximum context for one request is 128K tokens. On a 24 GB card that needs a quantized KV cache (f16 loads but cannot serve); the cache type is part of the serving profile and is declared here rather than varied per host. The row below was measured with a 128K window configured and a short prompt, so it shows the window fits and serves, not decode speed with a fully populated 128K context, which has not yet been measured on this hardware:
- Configured window / KV cache
- 131,072 · q8_0
- Decode, client-side (short prompt)
- 59.66 tok/s
- Decode, server-reported (short prompt)
- 76.01 tok/s
- VRAM
- 20,546 MiB
What the columns mean
- Single-request decode — output tokens per second for one request with no other load present. The number a client experiences.
- Aggregate — total output tokens per second across all active streams. This is what determines cost per token.
- Per stream — output tokens per second for one request once the GPU is shared across N streams. Falls as concurrency rises.
- Prompt processing — prefill throughput, counted separately because it is not shared across streams the way decode is.
- Serving cost / M out — derived from the measured host hourly rate divided by measured aggregate output. The cost basis for a published price, never the price itself.
Evidence grades
Every figure on this site carries one of these. They are deliberately styled differently so a projection never reads as a measurement.
- MEASURED Observed on the serving configuration above.
- DERIVED Calculated from measured inputs; the arithmetic is published.
- ESTIMATED Projected from related measurements, not directly observed.
- CURRENT MARKET An external figure that changes, re-fetched when used.
Published figures are cross-checked against their source benchmark at build time. If a re-run invalidates a number here, the build fails rather than the stale figure staying online.