gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Draft model
incoai/Qwen3.8-27B-DFlash2
Workload
1024 → 256
Runs
9 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)01343.904run_aI4FkZvVY0p1KI45: 4.4 req/s, p95 TTFT 1,343.9 msrun_okGtXqxwISTiqbue: 4.4 req/s, p95 TTFT 225.2 msrun_lCtZaPCT6by0uscg: 4.4 req/s, p95 TTFT 224.7 msrun_lVl-eHrXjQFFtVVq: 4.39 req/s, p95 TTFT 225 msrun_HdWZbZ-8Ez77f3Uu: 2.61 req/s, p95 TTFT 136.2 msrun_mATdf5hwQdeooXoO: 1.49 req/s, p95 TTFT 83.5 msrun_QPTe584j4T2arbz0: 1.51 req/s, p95 TTFT 83.2 msrun_T7Riy9CpQJph9INq: 1.51 req/s, p95 TTFT 83.2 msrun_kwpo7O9vGCtX25x8: 1.51 req/s, p95 TTFT 83.1 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=81,126.51,343.9 ms3.5 ms0run_aI4FkZvVY0p1KI45
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,127.4225.2 ms3.5 ms0run_okGtXqxwISTiqbue
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,125.2224.7 ms3.5 ms0run_lCtZaPCT6by0uscg
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=41,124.1225 ms3.71 ms0run_lVl-eHrXjQFFtVVq
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=2668.1136.2 ms2.88 ms0run_HdWZbZ-8Ez77f3Uu
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4c=1382.583.5 ms2.68 ms0run_mATdf5hwQdeooXoO
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1385.983.2 ms2.66 ms0run_QPTe584j4T2arbz0
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=1385.983.2 ms2.66 ms0run_T7Riy9CpQJph9INq
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=138683.1 ms2.66 ms0run_kwpo7O9vGCtX25x8
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=8
Req/s
Output tok/s
1,126.5
p95 TTFT / TPOT
1,343.9 ms / 3.5 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,127.4
p95 TTFT / TPOT
225.2 ms / 3.5 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,125.2
p95 TTFT / TPOT
224.7 ms / 3.5 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=4
Req/s
Output tok/s
1,124.1
p95 TTFT / TPOT
225 ms / 3.71 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=2
Req/s
Output tok/s
668.1
p95 TTFT / TPOT
136.2 ms / 2.88 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=4 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=4
Test load
c=1
Req/s
Output tok/s
382.5
p95 TTFT / TPOT
83.5 ms / 2.68 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
385.9
p95 TTFT / TPOT
83.2 ms / 2.66 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
385.9
p95 TTFT / TPOT
83.2 ms / 2.66 ms
Runtime
sglang 0.0.0.dev1+g5f55db35eDFlash2 · BF16 draft
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
386
p95 TTFT / TPOT
83.1 ms / 2.66 ms