Benchmarks

Measurements, with the configuration attached.

A throughput number without the hardware, the build and the load it was taken under is not a measurement. Everything below is reported with the conditions that produced it.

Single request

What one un-contended request decodes at, and prompt processing. This is the figure a client experiences before any other load is present.

Single-request throughput and prompt processing
ModelSingle-request decodePrompt processingBenchmark contextGrade
Qwen3.8-27B63.70 tok/s1130 tok/s4,096 ctxMEASURED

Decode vs concurrency

All streams fired simultaneously, 256 completion tokens each. Aggregate is total output across every active stream; per stream is what one of them experiences once the GPU is shared. Measured at 4,096 context with an f16 KV cache and short prompts — concurrency with long live contexts has not been measured, and a long context both takes KV memory and slows each decode step.

Decode throughput across concurrency levels
ModelStreamsAggregatePer streamVRAMServing cost / M out
Qwen3.8-27B163.68 tok/s63.70 tok/s15,298 MiB$0.779
281.19 tok/s40.83 tok/s15,892 MiB$0.611
3114.22 tok/s39.03 tok/s16,528 MiB$0.434
4127.19 tok/s32.47 tok/s17,102 MiB$0.390
8167.20 tok/s21.73 tok/s19,550 MiB$0.297

Aggregate does not scale linearly with concurrency: measured throughput reached 2.6× at eight streams, not 8×. Per stream therefore falls as concurrency rises, and the two columns answer different questions — one is capacity, the other is the latency a user feels. Serving cost is derived from the measured host rate and aggregate throughput; it is the cost basis a customer price is set from, not a published price.

Serving configuration

The build and the flags are ours and are recorded, because a host choosing them would make every number above unreproducible.

GPU
NVIDIA GeForce RTX 3090 (24,576 MiB)
Host
Vast.ai (verified) · US / SE (two runs) · $0.1786/hr
Runtime
llama.cpp 0.5.0-dev (build 11235, commit 6c7a87f7e)
Driver
595.71.05, CUDA 13.2
Batch / ubatch
2048 / 512

Artifact

The exact weights served, identified down to the revision hash.

Quant
UD-IQ4_XS
Size
14.25 GB
Upstream
unsloth/Qwen3.8-27B-GGUF
Revision (sha256)
40fac4050e940397dbf13087afd50f4734a11805bf9d65ef8ddd7483470e6199
License
apache-2.0

Context and KV

The maximum context for one request is 128K tokens. On a 24 GB card that needs a quantized KV cache (f16 loads but cannot serve); the cache type is part of the serving profile and is declared here rather than varied per host. The row below was measured with a 128K window configured and a short prompt, so it shows the window fits and serves, not decode speed with a fully populated 128K context, which has not yet been measured on this hardware:

Configured window / KV cache
131,072 · q8_0
Decode, client-side (short prompt)
59.66 tok/s
Decode, server-reported (short prompt)
76.01 tok/s
VRAM
20,546 MiB

What the columns mean

  • Single-request decode — output tokens per second for one request with no other load present. The number a client experiences.
  • Aggregate — total output tokens per second across all active streams. This is what determines cost per token.
  • Per stream — output tokens per second for one request once the GPU is shared across N streams. Falls as concurrency rises.
  • Prompt processing — prefill throughput, counted separately because it is not shared across streams the way decode is.
  • Serving cost / M out — derived from the measured host hourly rate divided by measured aggregate output. The cost basis for a published price, never the price itself.

Evidence grades

Every figure on this site carries one of these. They are deliberately styled differently so a projection never reads as a measurement.

  • MEASURED Observed on the serving configuration above.
  • DERIVED Calculated from measured inputs; the arithmetic is published.
  • ESTIMATED Projected from related measurements, not directly observed.
  • CURRENT MARKET An external figure that changes, re-fetched when used.

Published figures are cross-checked against their source benchmark at build time. If a re-run invalidates a number here, the build fails rather than the stale figure staying online.