gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
maurienne-ai/Qwen3.8-27B-DFlash2-NVFP4-RTNcal
Workload
1024 → 256
Runs
9 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)0541.106run_Csl1wzmwaa5WbXqq: 6.17 req/s, p95 TTFT 541.1 msrun_wffL2aOab4uxofSZ: 6.11 req/s, p95 TTFT 525.7 msrun_Ua9LfctoL62V1zhZ: 6.05 req/s, p95 TTFT 531.6 msrun_TVD-yEbfI7znExhm: 4.63 req/s, p95 TTFT 163.2 msrun_tQ08GsRht9cMuN-4: 2.81 req/s, p95 TTFT 147.3 msrun_F4Kz3F-9yzkw5N-l: 1.62 req/s, p95 TTFT 81.8 msrun_SEz8DAmLekmNP5AB: 1.62 req/s, p95 TTFT 81.5 msrun_qX0tR01cWoRMSr42: 1.62 req/s, p95 TTFT 81.5 msrun_nmXMEBpM2V5devtc: 1.62 req/s, p95 TTFT 81.7 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=81,578.6541.1 ms4.75 ms0run_Csl1wzmwaa5WbXqq
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=81,563.2525.7 ms5.04 ms0run_wffL2aOab4uxofSZ
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=81,549.7531.6 ms5.02 ms0run_Ua9LfctoL62V1zhZ
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=41,185.6163.2 ms3.42 ms0run_TVD-yEbfI7znExhm
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2718.4147.3 ms3.03 ms0run_tQ08GsRht9cMuN-4
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=1414.781.8 ms2.56 ms0run_F4Kz3F-9yzkw5N-l
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1415.481.5 ms2.56 ms0run_SEz8DAmLekmNP5AB
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1415.581.5 ms2.56 ms0run_qX0tR01cWoRMSr42
sglang 0.0.0.dev1+g5f55db35eDFlash2 · NVFP4dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1414.681.7 ms2.57 ms0run_nmXMEBpM2V5devtc
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
1,578.6
p95 TTFT / TPOT
541.1 ms / 4.75 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
1,563.2
p95 TTFT / TPOT
525.7 ms / 5.04 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
1,549.7
p95 TTFT / TPOT
531.6 ms / 5.02 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
1,185.6
p95 TTFT / TPOT
163.2 ms / 3.42 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
718.4
p95 TTFT / TPOT
147.3 ms / 3.03 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
414.7
p95 TTFT / TPOT
81.8 ms / 2.56 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
415.4
p95 TTFT / TPOT
81.5 ms / 2.56 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
415.5
p95 TTFT / TPOT
81.5 ms / 2.56 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · calibrated NVFP4 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
414.6
p95 TTFT / TPOT
81.7 ms / 2.57 ms