gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090

Hardware
1× NVIDIA GeForce RTX 5090
Workload
1024 → 256
Runs
9 eligible
Efficient frontierOther eligible run
p95 TTFT (ms)Request throughput (req/s)0501.402run_Rhgr51rIpq9Jzp6l: 2.22 req/s, p95 TTFT 500.7 msrun_Z6WOxGffTnaDOo_i: 2.22 req/s, p95 TTFT 500.7 msrun_3yAiq2C-hBCAw8Xo: 2.22 req/s, p95 TTFT 501.4 msrun_6YNsQIR1e2J7DaLI: 1.16 req/s, p95 TTFT 255.8 msrun_wJicBVkg42sBm-Go: 0.62 req/s, p95 TTFT 133.7 msrun_SysUk70j-HqYOBbo: 0.34 req/s, p95 TTFT 72.9 msrun_Vud0kxXwRLFdF-bd: 0.34 req/s, p95 TTFT 72.9 msrun_DTV6eM8cWgnnQS9C: 0.34 req/s, p95 TTFT 72.9 msrun_j9dHG1N-lX_-hBZA: 0.34 req/s, p95 TTFT 73 ms

Runs

All published runs with the same model build, accelerator setup, and request shape. Benchmark methodology remains attached to each run.

RuntimeMaterial parametersTest loadReq/sOutput tok/sp95 TTFTp95 TPOTp95 E2EFailuresRun
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8569.4500.7 ms13.6 ms0run_Rhgr51rIpq9Jzp6l
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8567.5500.7 ms13.6 ms0run_Z6WOxGffTnaDOo_i
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=8567.6501.4 ms13.6 ms0run_3yAiq2C-hBCAw8Xo
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=4297.3255.8 ms13.01 ms0run_6YNsQIR1e2J7DaLI
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=2158.2133.7 ms12.19 ms0run_wJicBVkg42sBm-Go
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8c=187.372.9 ms11.22 ms0run_SysUk70j-HqYOBbo
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=187.772.9 ms11.18 ms0run_Vud0kxXwRLFdF-bd
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=187.872.9 ms11.18 ms0run_DTV6eM8cWgnnQS9C
sglang 0.0.0.dev1+g5f55db35edtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1c=187.473 ms11.22 ms0run_j9dHG1N-lX_-hBZA
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
569.4
p95 TTFT / TPOT
500.7 ms / 13.6 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
567.5
p95 TTFT / TPOT
500.7 ms / 13.6 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=8
Req/s
Output tok/s
567.6
p95 TTFT / TPOT
501.4 ms / 13.6 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=4
Req/s
Output tok/s
297.3
p95 TTFT / TPOT
255.8 ms / 13.01 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=2
Req/s
Output tok/s
158.2
p95 TTFT / TPOT
133.7 ms / 12.19 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=8 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=8
Test load
c=1
Req/s
Output tok/s
87.3
p95 TTFT / TPOT
72.9 ms / 11.22 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
87.7
p95 TTFT / TPOT
72.9 ms / 11.18 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
87.8
p95 TTFT / TPOT
72.9 ms / 11.18 ms
Runtime
sglang 0.0.0.dev1+g5f55db35e
Material parameters
dtype=bfloat16 context_length=9216 kv_cache_dtype=fp8_e4m3 attention_backend=flashinfer cuda_graph_max_bs=1 mem_fraction_static=0.9 chunked_prefill_size=1024 max_running_requests=1
Test load
c=1
Req/s
Output tok/s
87.4
p95 TTFT / TPOT
73 ms / 11.22 ms