← 开源
1CatAI

1Cat-vLLM

V100 / SM70-focused vLLM engineering fork for modern LLM inference.

InfrastructureKernelsInference service optimizationPython
在 GitHub 打开
增长势头
+924 小时新增 Star+0.7%
1.23k
Star
242
Fork
+53
本周
100
贡献者
创建于 2026-03-10 · 更新于 2026-10-05 · 今日第 1026 名
主要开发者
README

1Cat-LLM logo

1Cat-LLM

Make Volta Fast Again

Modern LLM inference for NVIDIA Tesla V100 / SM70

1Cat-vLLM is becoming 1Cat-LLM.

Broader support. More efficient kernels. More models to choose from.

We started by making modern models fast on V100. The next chapter carries that work further: broader model and runtime support, more efficient execution, and more choice in how models are packaged and quantized.

New GGUF support brings a wider choice of community-quantized checkpoints for supported architectures, with more options for model size, memory use and quality. The new name reflects that direction, with Volta engineering at its foundation.

The rename is planned. Repository links, release filenames and installation commands below retain their current names during the transition.

Recommended checkpoints:

Target: QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4

MoE: RadixArk/Qwen3.8-Flash-Next-NVFP4

DFlash2 draft: incoai/Qwen3.8-27B-DFlash2

4× Tesla V100 16GB · Qwen3.8-27B-NVFP4 + DFlash2 · ≈260 tok/s

Tesla V100 was released in 2017.

Its Tensor Cores did not suddenly become useless.

The software stack simply stopped being optimized seriously for SM70.

1Cat-LLM grows out of 1Cat-vLLM, a vLLM engineering fork that treats NVIDIA Volta / SM70 / Tesla V100 as a first-class optimization target.

We are not satisfied with:

“The latest model can start on V100.”

Our goal is:

Make modern models actually run fast on V100.

The recorded four-card Tesla V100 16GB demo runs Qwen3.8-27B-NVFP4 + DFlash2 through 1Cat-vLLM at roughly:

≈260 tokens/s

Demo: 4× V100 running Qwen3.8-27B-NVFP4-DFlash2

≈260 tok/s is a real-machine demo headline, not a universal fixed decode rate.

Every benchmark below retains its own hardware, model, context length, batch size, KV dtype, sampling policy, and speculative-decoding contract. Attention TFLOP/s, prefill tok/s, target-only decode tok/s, and speculative decode tok/s are not interchangeable metrics.

Latest published release: v1.5.1, October 3, 2026. The release measurements below use 4× V100-SXM2 32GB; they are a separate contract from the 16GB demo.

This README also covers merged development work through main b379775d, October 5, 2026. The newer GGUF and mixed-load changes require a corresponding source/development build; they are not included merely by installing the v1.5.1 release wheel.


📊 Performance First

SM70 Flash-V100 resolves --kv-cache-dtype fp8 to E4M3; the 27B release profile explicitly selects fp8_e4m3. DFlash2 verification retains FP32 attention state and FP32 logits on its admitted paths.

Keep the Python package and Flash-V100 extension from the same build. The original E4M3 repair requires precision revision 4; current multi-head and request-major routes require newer native capabilities. See the precision contract and concurrent long-context validation. Historical E5M2/FP16-partial performance results below keep their original configuration and are not speed claims for these precision defaults.

v1.5.1: Measured Release-to-Release Decode

Input / workload v1.5.0 v1.5.1 Comparison
4K · single request 103.87 tok/s 229.34 tok/s +120.80%
32K · single request 49.04 tok/s 201.18 tok/s +310.24%
128K · single request No completed baseline 171.43 tok/s No speedup calculated
32K · C4 common decode window Not measured under the same contract 445.45 aggregate tok/s New-version result only

Speedup percentages are calculated from the displayed throughput values.

Single-request contract: official wheels, the same 4× V100-SXM2-32GB / TP4, QUASAR 27B NVFP4 target, DFlash2 q7, FP16 compute/draft KV, E4M3 target KV, CUDA 12.8 / Torch 2.10, prefix caching, asynchronous scheduling and CUDA Graphs. Both versions use the 1.5.1 launch settings: 262144 maximum context, 8192-token prefill budget, four sequence slots, memory utilization 0.80, KV/Mamba blocks 2048/8192. Sampling is T=1, top-p=.95, top-k=20, seed=0, up to 256 output tokens with normal EOS. Each value is the median of three warm requests after one cold request, from one successful startup per version.

Pure decode excludes TTFT/prefill. At 32K, the old wheel reused 28,672 prefix tokens while the new wheel recomputed the prompt; both decoded with the same 32,768-token input context. This is a comparison under one shared deployment recipe, not each version's best possible tuning. The old 128K request did not finish within the 180-second test timeout.

C4 contract: fixed 32K synthetic inputs, 256 forced output tokens per request, T=.7 / top-p=.8 / top-k=20, fixed per-request seeds, one warmup cohort and three measured cohorts from one startup. Throughput counts returned token IDs only while all four requests are decoding. Forced-length timing is separate from natural-EOS quality testing.

Evidence and complete settings: v1.5.1 release notes. These release checks include output-health spot checks, not a new broad dataset-quality campaign.

Long-Context Attention: 17.92 → 47.1 → ≈60.8 → ≈71 TFLOP/s

Stage Evidence Useful causal Attention compute Notes
Previous production path v1.2.2-era baseline 17.92 TFLOP/s V100 long-prefix Attention baseline
D256 Split-D / N32 v1.3.0 46.63–47.1 TFLOP/s ≈2.6× over the previous production path
GQA-packed wide QK/PV PR #286 · historical measurement ≈60.8 TFLOP/s 6 GQA heads packed into wider Tensor-Core GEMMs
FP32 QK/PV · wider local head layouts PR #666 · Q8192/KV262144 69.982–71.111 TFLOP/s Single-V100 operator replay for TP1/TP2/TP4 local layouts, including head-group copies
Earlier experimental ceiling PR #315 ≈79 TFLOP/s Historical candidate; not a Release/default quality claim

From 17.92 → ≈60.8 TFLOP/s, representative long-context V100 Attention useful compute improved by roughly 3.4× on the same generation of hardware.

These figures count useful causal QK/PV work, not whole-model TOPS.

The later FP32 QK/PV record is 71.111 / 70.221 / 69.982 TFLOP/s for TP4 / TP2 / TP1 local head layouts. These are operator measurements on one V100-SXM2-32GB, not multi-GPU model throughput. Earlier 75–77T records predate the complete QK FP32 repair; they are not the current precision baseline. Power and clock policy also matter: the same-source 185W audit recorded 53.816T. See shape coverage and quality limits.


🚀 Real Model Benchmarks

The table below preserves complete-model / API / pure-decode / speculative-decode measurements from their original PRs. They are historical evidence under separate contracts; the current published-wheel comparison is above.

Model Hardware / Runtime Workload Measured result Evidence / Status
Qwen3.6-27B-AWQ + MTP4 4× V100 · TP4 · E5M2 KV · Flash-V100 · CUDA Graph 64K decode 100.564 tok/s v1.2.2 Release · AL 4.981 / 99.52%
Qwen3.6-27B-AWQ + MTP4 same 128K decode 85.258 tok/s v1.2.2 Release · +87.64% vs no-MTP
Qwen3.6-27B-AWQ + MTP4 same · max 256K 261,888 context decode 49.772 tok/s v1.2.2 Release · AL 5.000 / 100%
Qwen3.6-35B-A3B NVFP4 4× V100 · TP4 · mixed FP8 + W4A16_NVFP4 4096 / 1024 · no-MTP 116.99 tok/s #270
Qwen3.6-35B-A3B NVFP4 + MTP4 same matched MTP4 run 174.76 tok/s #270 · 1.49× no-MTP
Qwen3.8-27B-NVFP4 4× V100 · TP4 · E4M3 KV · full CUDA Graph · no-MTP exact 128K decode 61.834 tok/s #285 · measured
Qwen3.8-27B-NVFP4 same exact 256K decode 50.376 tok/s #285 · measured, not projected
Qwen3.8-27B-FP8 4× V100 · TP4 · E5M2 KV · no-MTP 128K decode 50.68 tok/s #212 release-path sweep
Qwen3.8-27B-FP8 same 256K decode 41.11 tok/s #212 release-path sweep
Qwen3.8 Flash-Next-NVFP4 4× V100 · TP4 · V2 · full CUDA Graph · no-MTP 8K / 512 pure decode 80.732 tok/s #415 · quality-audited
Qwen3.8 Flash-Next-NVFP4 + MTP4 4× V100 · TP4 · V2 final cold-JIT gate 138.26 tok/s #389 · AL 4.943 / 98.57%
Qwen3.8-27B-NVFP4 + DFlash2 4× V100 · TP4 · production API historical web prompt · 512 output 206.06 tok/s streaming decode #422 · 17.463 ms/round · 3.599 emitted/round
Qwen3.8-27B-NVFP4 + DFlash2 4× V100 · TP4 · practical API MBPP item 28 · natural EOS 251.60 tok/s #288 · AL 4.686 · EvalPlus 1/1
Qwen3.8 DFlash2 + adaptive lookup q16 4× V100 · TP4 · opt-in lookup augmentation repeated-context sample 316.27 tok/s #366 · 3.162 ms TPOT · special opt-in contract
DeepSeek-V4-Flash 8× V100 · TP8 · FP8 dense + MXFP4 experts · CUDA Graph · no-spec 1024 / 256 15.357 ms TPOT ≈ 65.1 tok/s #181 · accepted no-MTP baseline
DeepSeek-V4-Flash 8× V100 · PP2×TP4 · no-DSpark combined quality-checked endpoint 73.613–73.646 tok/s #344
DeepSeek-V4-Flash same PP2×TP4 strict control dataset-quality pair 73.539 tok/s #344 · GSM8K 64/64 · HumanEval 29/32
GLM-5.3-Flash-NVFP4 8× V100 · TP4/PP2 · E4M3 KV · no-MTP 1K / 256 decode 53.016 tok/s #402 · merged historical quality audit

🧪 Dataset / Quality × Throughput Benchmarks

Raw tok/s alone can turn optimization into a benchmark game. 1Cat-LLM therefore records real model throughput, dataset score, natural-stop health, output validity, and speculative acceptance together.

The retained DFlash2 coding/PPL gates below use their recorded historical configuration, including E5M2 KV where specified. They do not stand in for a new quality evaluation of the E4M3 release wheel, current main, or GGUF models.

Qwen3.8-27B-NVFP4 + DFlash2 — Practical 16K coding gate

Contract:

  • 4× V100, TP4
  • NVFP4 target
  • official BF16 DFlash2 drafter
  • FP8 E5M2 target KV
  • FlashAttention-V100
  • full CUDA Graph
  • prefix cache
  • Mamba align
  • temperature=1.0
  • top_p=0.95
  • top_k=20
  • xhigh reasoning
  • 16K natural-EOS output cap
  • three predeclared sampling seeds
Dataset Samples Base score Plus score Natural stop Aggregate output throughput Mean steady decode Acceptance pooled / request
MBPP / EvalPlus 96 · 93 scored 89/93 80/93 95/96 213.539 tok/s 236.902 tok/s 4.061 / 4.318
HumanEval / EvalPlus 96 94/96 92/96 91/96 208.978 tok/s 245.645 tok/s 3.972 / 4.476

Evidence: PR #346 and docs/design/sm70_dflash2_quality_audit.md.

The first seed matches the historical target-only / no-DFlash request-seed contract. Across MBPP + HumanEval, both routes score:

Base : 62 / 63
Plus : 59 / 63

So the 200+ tok/s speculative path is not obtained by removing the task-quality gate.

About length-capped failures

Some coding failures are caused by long reasoning exhausting the 16K output budget, rather than by an invalid final solution.

Across the retained MBPP + HumanEval campaign, 6 of 192 outputs reached the 16K cap, and 3 of those were still extractable and correct.

For this reason, the README separates:

  • executable task score,
  • natural-stop rate,
  • output-cap failures,
  • and throughput.

It does not treat every capped sample as proof of a model-capability regression.

Optional precise-coding profile

Using the same DFlash2 engine and the same middle seed, only client-side sampling is changed to the model's precise-coding profile:

temperature = 0.6
top_p       = 0.95
top_k       = 20
Dataset Temperature 1.0 Temperature 0.6
MBPP Base 27/31 · Plus 24/31 Base 29/31 · Plus 27/31
HumanEval Base 31/32 · Plus 31/32 Base 31/32 · Plus 31/32

Across 80 requests:

Mean steady decode:
233.187 → 244.520 tok/s

Output-token / decode-time throughput:
195.817 → 201.852 tok/s

Request-mean acceptance:
4.27345 → 4.51356

Natural stops move from 72/80 to 70/80, so this remains an optional precise-coding profile, not a forced global default.


Other full-model quality gates

Model / Route Dataset / Quality Real throughput under the recorded contract Status
Qwen3.8 Flash-Next-NVFP4 · no-MTP GSM8K 15/16 raw · 15/16 strict · 16/16 natural stop 80.935 tok/s weighted pure decode #415 · merged / quality-audited
Qwen3.8 Flash-Next-NVFP4 · MTP4 HumanEval8 8/8 semantic executions 150.17 tok/s weighted pure decode #398 · merged; bounded historical gate
Qwen3.6-35B-A3B NVFP4 + MTP4 GSM8K 122/128 (95.3125%) · 0 invalid · 0 repetitive matched MTP run 174.76 tok/s #270 · merged
Qwen3.6-35B-A3B NVFP4 + MTP4 ShareGPT16 final-SHA workload 120.096 tok/s pure decode · 97.678 E2E output tok/s · 241.973 prefill tok/s #270 · merged
DeepSeek-V4-Flash · PP2×TP4 GSM8K 64/64 · HumanEval 29/32 · LongBench 44.740 73.539 tok/s median #344 · strict quality control
DeepSeek-V4-Flash · PP2×TP4 Combined route endpoint: GSM8K 62/64 · 0 invalid · coherent output 73.613–73.646 tok/s #344
GLM-5.3-Flash-NVFP4 · no-MTP Max reasoning: 6/8 tasks finish within 4096 output tokens; targeted low-reasoning code rerun 2/2 AST + execution 53.016 tok/s decode · 266.040 tok/s 1K prefill #402 · merged; historical quality matrix

Tool Calling / Structured Output

Modern inference serving must do more than generate prose. DFlash2 + adaptive lookup was also tested against tool and structured-output workloads.

Gate Result
BFCL 29/32
ToolACE 12/12
NexusRaven 13/16
Strict JSON Schema 7/8
Structured B1 12/12
Structured B4 12/12
Long prefix-state isolation 5/5

These quality results match the target-only / q7 reference in the retained audit.

Runtime examples from the same development line:

Ordinary q7 → adaptive q8:
168.52 → 170.98 tok/s

Repeated-context q16:
316.27 tok/s
3.162 ms TPOT

The q16 number is a special repeated-context lookup-hit contract. It is not presented as the expected throughput of every tool-calling request.

Evidence: PR #366.


Distribution / PPL gate

DFlash2 is also checked at the target-distribution level.

Eight fixed WikiText 2,048-token segments, 16,376 scored prompt tokens:

Target-only PPL : 5.4993116
DFlash2 PPL     : 5.4993622
Absolute delta  : +0.0000506
Relative delta  : +0.00092%
Max segment Δ   : 0.0062143

The purpose of this gate is to detect cases where benchmark answers still look acceptable while speculative verification has systematically shifted the target distribution.


⚡ Why DFlash2 Can Reach 200+ tok/s

The repository contains multiple real full-model DFlash2 throughput records:

  • production web prompt: 206.06 tok/s streaming decode, 512 output tokens, 17.463 ms/engine round;
  • high-acceptance MBPP request: 251.60 tok/s, acceptance length 4.686, 328-token natural EOS, EvalPlus Base/Plus 1/1;
  • adaptive lookup q16 repeated-context workload: 316.27 tok/s, explicitly a special opt-in repeated-context contract;
  • the README headline remains ≈260 tok/s from the real-machine demo.

206, 251, 260, and 316 tok/s are not the same benchmark.

DFlash2 throughput depends strongly on acceptance length, prompt repetition, q8/q16 verification width, context length, and task type.


📏 Long Context Means More Than “It Fits in 256K”

For Qwen3.8-27B-NVFP4, PR #285 reports real TP4 full-model long-context decode:

128K : 40.561 → 61.834 tok/s
256K : 27.456 → 50.376 tok/s

50.376 tok/s at 256K is the measured endpoint result.

The PR also contains a decomposition-based projection of 52.216 tok/s, but this README intentionally uses the measured 50.376 tok/s result.

DeepSeek-V4 should also be judged by later full-model results rather than an early bring-up checkpoint:

TP8 no-spec:
15.357 ms TPOT ≈ 65.1 tok/s

PP2×TP4 quality-checked endpoint:
73.613–73.646 tok/s

🔬 Selected Merged PR Benchmarks

Area PR / Contract Control 1Cat result Gain
D256 long-prefill Attention #198 · Q4096/KV64K · Hq6/Hkv1/D256 87.6001 ms 50.4504 ms 1.74×
D256 long-prefill Attention #198 · Q4096/KV8K 11.1255 ms 5.0542 ms 2.20×
128-bit E5M2 XQA load #268 · B16/17.8K operator 0.743424 ms 0.602112 ms 1.235×
128-bit E5M2 XQA load #268 · ragged B16/32K operator 1.171296 ms 0.925808 ms 1.265×
Batched long decode #268 · B16/16K full-model pure decode 529.071 tok/s 570.982 tok/s +7.92%
Long-context decode routing #206 · 128K TP4 40.8208 tok/s 48.5431 tok/s +18.92%
Long-context decode routing #206 · 180K TP4 36.1387 tok/s 42.5501 tok/s +17.74%
E4M3 XQA long decode #285 · exact 128K 40.561 tok/s 61.834 tok/s +52.45%
E4M3 XQA long decode #285 · exact 256K 27.456 tok/s 50.376 tok/s +83.48%
Grouped QSA Page4 #387 · per-layer/rank 55.151 ms 9.632 ms 5.518×
QSA full-model prefill #387 · 64K 4,446.64 tok/s 5,777.43 tok/s +29.93%
Indexed NVFP4 MoE prefill #390 · 64K 5,777.43 tok/s 6,241.48 tok/s +8.03%
Exact target-only decode #415 · 8K/512 · no-MTP 65.864 tok/s 80.732 tok/s +22.57%
DFlash2 NVFP4 prefill #417 · 32K/64K retained pre-closure 4069.25 / 3566.94 prefill tok/s +30.1% / +37.7%
DeepSeek-V4 sparse MLA #163 · sparse MLA GPU service 46.920 ms/token 4.392 ms/token -90.64%
DeepSeek-V4 TP8 no-spec decode #181 · 8×V100 · 1024/256 19.342 ms TPOT false-4K graph 15.357 ms TPOT ≈65.1 tok/s ~20.6% lower TPOT
DeepSeek-V4 PP2×TP4 full model #344 · 8×V100 · no-DSpark — 73.613–73.646 tok/s quality-checked endpoint
DFlash2 concurrent long decode #697 · 32K/256 · C4 / C8 326.832 / 492.025 tok/s 429.134 / 624.863 tok/s +31.30% / +27.00%
Flash-Next prefix-cache prefill #754 · 32K · zero prefix hits 3,046 tok/s 5,068 tok/s +66.38%
Flash-Next NVFP4 MTP4 defaults #796 · 31,744/256 · C1 79.832 tok/s 92.199 tok/s +15.49%
Original-byte IQ3_XXS/IQ3_S pair #966 · M8/N4352/K5120 · operator only 86.016 μs 57.344 μs 1.50×

The added rows retain their own baselines. #697 uses 4× V100-SXM2-32GB / TP4, NVFP4 + DFlash2 q7, E4M3 target KV and FP16 draft KV, with three warm simultaneous-request decode windows; acceptance also changes. #754 uses Flash-Next NVFP4 / MTP4, FP16 KV, prefix caching enabled and matched no-hit requests. #796 uses the same FP16-KV MTP4 recipe in both arms, with 21/21 exact paired quality outputs; its C4 change is only +1.05%, and additional packed weights cost about 1.25 GiB/rank. #966 is a fixed-clock single-GPU cold-L2 operator comparison, not a whole-model speedup. These gains cannot be added together or substituted for the release-wheel comparison.


🧠 128-bit Loads: Not a Cosmetic Vectorization Change

PR #268 does more than replace a narrow type with a wider C++ type.

Inside real paged-KV partitions, it:

  • reuses the Page ID;
  • merges two half8 conversion groups;
  • issues one aligned 128-bit cache load;
  • keeps softmax, PV, partition boundaries, and reduction order unchanged.

NCU evidence:

L1 global-load requests:
656,443 → 383,814
-41.53%

Executed warp instructions:
97,998,831 → 83,583,696
-14.71%

Long-scoreboard stall:
39.14% → 30.10%

Eligible warps / scheduler:
0.55 → 0.65

B16 / 17.8K kernel duration:
648.352 → 499.520 μs
-22.95%

DRAM bytes stay nearly unchanged.

The gain comes from fewer fragmented loads, lower address/dependency pressure, and a more continuous operand feed, not from magically reducing the model size.


✅ Correctness / Quality Gates

1Cat-LLM does not treat a good-looking TPS number as sufficient evidence.

Representative gates include:

  • #198: 64K full-model A/B/A 64-token IDs, text, and SHA256 match; random paged-KV, gathered-dense, and Split-KV3 have separate numerical gates.
  • #268: uniform/ragged B4/B8/B12/B16, page256/page800, 12K–32K operator A/B is bitwise exact.
  • #285: 128K and 256K E4M3 XQA endpoints both emit the complete 64 tokens and preserve their matching control streams.
  • #346: structured API 24/24, long alternating-prefix state 5/5, multi-seed MBPP/HumanEval quality gates, and target-only/DFlash2 WikiText PPL 5.4993116 / 5.4993622.
  • #387: grouped QSA replay is deterministic; arithmetic, Chinese-language, and performance-case token hashes match the retained baseline.
  • #415: GSM8K 15/16 strict, natural stop 16/16, zero capped outputs, zero structurally invalid outputs.
  • #427: 1.5.0 RC isolated install passes /v1/models, /metrics, normal chat, streaming/non-streaming tool calls, JSON Schema, and repeated-prefix checks; a 10,017-token prefix moves from 2.642 s cold → 0.164 s cached.

🔥 FlashAttention-V100

We are not just “making FlashAttention compile on V100.”

We are rebuilding the dataflow for Volta

FlashAttention is fundamentally an IO and scheduling problem:

  • reduce HBM round trips;
  • keep Q/K/V and intermediate state on-chip as long as possible;
  • increase reuse;
  • reduce materialization;
  • reduce barriers;
  • continuously feed Tensor Cores.

Modern FlashAttention implementations are designed around Ampere, Hopper, and newer GPUs.

Tesla V100 is SM70.

It does not have:

  • Ampere cp.async;
  • Turing/Ampere-style ldmatrix data paths available to newer Tensor-Core kernels;
  • Hopper TMA;
  • native FP8 Tensor Cores;
  • Blackwell FP4 Tensor Cores.

A direct compatibility port may run, but it often leaves the GPU underfed.

That is why 1Cat-LLM rebuilds the execution path around the capabilities Volta actually has.


⚙️ Software-Reconstructed Async / Matrix Feed on SM70

We do not claim that V100 executes cp.async or ldmatrix.

Instead, 1Cat-LLM reconstructs the design goals behind those mechanisms using:

LDG
STS
LDS
register prefetch
double buffering
Shared Memory swizzle
explicit HMMA fragment mapping
cross-tile / cross-stage software pipelining

The objective is the same:

overlap memory movement with compute
        ↓
increase on-chip reuse
        ↓
shorten dependency chains
        ↓
reduce barriers and replay
        ↓
keep HMMA continuously fed

Representative techniques include:

  • register prefetch and double buffering;
  • overlap next-K tile loading with current QK compute;
  • pre-stage PV operands while HMMA is still executing;
  • phase-swizzled Shared Memory layouts;
  • 128-bit vectorized access;
  • explicit QK/TN and PV/TT HMMA fragment ownership;
  • software scheduling across tile and stage boundaries.

We do not emulate a cp.async instruction.

We rebuild the memory/computation overlap that modern hardware instructions are designed to provide.


Layer 1 — Move KV Cache Correctly, Wide, and Once

Paged KV maps logical tokens onto physical pages.

A naive SM70 path repeatedly:

load Page ID
calculate address
load narrow FP8 fragment
convert
repeat

That wastes cycles on address work, dependency waits, and scalar memory traffic.

The 128-bit XQA work in #268 reuses page metadata and performs paired aligned loads.

Representative full-model batch results include:

B16 / 16K:
529.071 → 570.982 tok/s
+7.92%

The corresponding operator gain reaches roughly 21%–26.5% on representative long-context XQA shapes.


Layer 2 — Rewrite D=256 Attention as a Volta-Native Pipeline

After reducing data-movement overhead, the Attention body itself is restructured.

Key components include:

D256 Split-D

Split D=256 into four D64 slices.

Paired warps share QK probability work while increasing PV parallelism without recomputing the same QK work.

N32 Online Softmax

Retain causal online-softmax and FP32 accumulation contracts without materializing a full score matrix.

K-stage Ping-Pong

Alternate K/D64 panels across Shared Memory stages to reduce barrier and wait pressure.

Split-KV3

Split long-prefix KV work into three partitions where useful, then merge FP32 partial state.

GQA Multi-Head Packing

Pack six GQA query heads into wider Tensor-Core work.

Wide QK / PV

Turn many fragmented small Tensor-Core operations into larger, more regular QK/PV GEMM-style work.

Prefix / Causal-Tail Separation

Schedule the fully visible long prefix separately from the exact causal tail and merge the online-softmax state.

This optimization family evolved through PR #198, later D256 / Split-KV3 work, v1.3.0, and PR #286.

The result:

17.92 TFLOP/s
    ↓
46.63–47.1 TFLOP/s
    ↓
≈60.8 TFLOP/s

Same GPU generation. Same Tensor Cores.

The software stopped wasting them.

≈79 TFLOP/s is retained as an experimental research ceiling, not as the default production quality claim.

Later work extends this dataflow to Q8000/Q8192, multiple requests and TP1/TP2/TP4 local head layouts, with FP32 QK/PV accumulation. The qualified full-FP32 operator record reaches roughly 70–71 TFLOP/s under the recorded conditions. Subsequent fixes cover short-chunk causal visibility, missed score maxima, overflowing score tiles and tail-intermediate range: #820, #850, #851, #875. Historical peak numbers do not replace these numerical checks.


Layer 3 — Sparse Attention Must Also Be Native to V100

Qwen3.8 Flash Next QSA requires more than “select fewer tokens.”

The runtime must also handle:

  • sparse block selection;
  • physical-page mapping;
  • Page4 K/V reuse;
  • exact per-row masks;
  • final QK/PV computation.

PR #387 groups eight adjacent query rows so overlapping Page4 K/V blocks are loaded once while preserving an exact 4-bit mask per row.

It then uses Volta WMMA directly for QK and PV.

Representative results:

Old QSA path:
55.151 ms/layer/rank

Grouped Page4:
9.632 ms/layer/rank
+0.362 ms planner

Attention speedup:
5.518×

Full-model pure-prefill improvements:

32K  : +32.36%
64K  : +29.93%
131K : +32.69%

🧩 Profiling-Driven Optimization

1Cat-LLM does not stop when one kernel becomes fast.

When QSA was accelerated, profiling showed the next hotspot had moved into NVFP4 MoE prefill.

PR #390 then removed the [tokens × topK, hidden] input-expansion bottleneck by using indexed W13 execution.

Representative results:

8K operator chain:
6.026752 → 4.235264 ms
1.423×

Full-model pure prefill:
32K  : 5998.65 → 6507.10 tok/s
64K  : 5777.43 → 6241.48 tok/s
131K : 5450.92 → 5871.47 tok/s

PR #393 then fused exact FP16 SwiGLU, split the N320 W13 tail into N256+N64, and removed wasted tail-tile work.

This is the optimization philosophy of the project:

Profile the real model, move the bottleneck, profile again.


🎯 Target-Only Decode Before Speculative Decoding

Before relying on DFlash2 or MTP, the target model itself must be fast.

PR #415 reports Qwen3.8-Flash-Next-NVFP4 on 4× V100:

8K input / 512 output
no MTP
full CUDA Graph

Control:
65.864 tok/s
15.183 ms TPOT

Candidate:
80.732 tok/s
12.387 ms TPOT

That is target-only throughput.

The route also passes:

GSM8K: 15/16 strict
Natural stop: 16/16
Weighted natural-output decode: 80.935 tok/s

⚡ DFlash2 on SM70

Traditional autoregressive decode requires one target-model pass per emitted token.

DFlash2 changes the execution model.

A block-diffusion draft model proposes several future tokens and the target verifies them together.

The effective service loop becomes:

draft several candidates
        ↓
target verifies a block
        ↓
accept multiple tokens
        ↓
advance by more than one token per target round

For Qwen3.8 DFlash2, the release-oriented SM70 stack also optimizes:

  • draft Attention;
  • selector;
  • grouped verifier;
  • GDN metadata;
  • sparse rejection;
  • NVFP4/QPN paths;
  • sampling;
  • CUDA Graph;
  • prefix state;
  • Mamba align;
  • tool / structured-output state.

The draft Attention itself uses:

FLASH_ATTN_V100

rather than falling back to an unrelated generic path.


DFlash2 Long-Context Decay

Long context must not make speculative verification cost grow unnecessarily.

PR #328 changes the non-anchored paged-prefill loop so it begins at the first sliding-window tile actually used by the draft.

At 256K:

Draft attention:
0.422912 → 0.246784 ms/layer

Five-layer projection:
2.114560 → 1.233920 ms

Candidate medians:

32K  : 0.252928 ms
128K : 0.243712 ms
256K : 0.246784 ms

The post-32K context slope is nearly eliminated for that draft-attention component.


🧵 Concurrency Must Survive New Prefill

Four decoders running alone are one workload. Four decoders sharing the GPU with new long prompts are another.

The newer SM70 path combines batched FP4/FP8 projections, joint draft attention, shared GDN metadata, long-context verification and sampling workspace reuse. #697, #706 and #708 retain separate operator and full-model comparisons.

PR #860 adds an adaptive mixed-prefill budget while preserving resident decode capacity and hybrid prefix replay. On its recorded 27B QUASAR NVFP4 + DFlash2 / 4× V100-SXM2-32GB / TP4 / E4M3 / 256K-capacity deployment, two sustained decoders share the GPU with two incoming 32K prompts:

Metric during incoming prefill Control Candidate
Resident output 6.88 tok/s 30.34 tok/s
Longest output-update gap 2.873 s 0.310 s
First / second incoming TTFT 14.287 / 20.631 s 17.994 / 35.491 s

These are medians of two warm measured runs after warmup, using the PR's two-NVLink-pair topology. Resident EOS is disabled to sustain the load; incoming requests complete naturally. Shorter decoder stalls trade against longer incoming TTFT. The 250 ms step target is a soft scheduler target, not a latency guarantee. This is merged development work after the v1.5.1 artifact.

Recent work also parallelizes compatible small-query draft attention (#944) and reduces synchronization in TP4 all-reduce plus RMSNorm (#957). Each keeps its measured precision, shape and topology limits; their small single-request gains are not advertised as a universal concurrency multiplier.


🔢 Quantization / Operator Stack

V100 predates many of the formats used by current LLM checkpoints.

1Cat-LLM therefore treats quantization support as an operator-design problem, not only a loader problem.

Current SM70 work includes:

  • AWQ / W4A16;
  • TurboMind SM70 kernels;
  • compressed-tensors;
  • FP8 E4M3 / E5M2 KV storage;
  • ModelOpt NVFP4;
  • MXFP4;
  • standalone GGUF metadata, tokenizer and architecture adapters;
  • original-block GGUF affine / codebook operators with capability-based fallback;
  • Quark W4A16 INT4 / UINT4;
  • QPN8;
  • QPN4;
  • QPN2;
  • grouped MoE;
  • exact-shape decode GEMV;
  • custom SM70 sampling paths.

The goal is not:

“The dtype parses.”

The goal is:

The quantized format becomes a usable high-performance serving path on Volta.


🗂️ GGUF: Keep the Checkpoint, Rebuild the Execution

More model choice starts with more usable checkpoints.

GGUF brings community quantizations into the supported model paths, giving users more ways to balance model size, memory use and quality. 1Cat-LLM connects that choice to native kernels and the serving stack, with coverage and measurements recorded below.

Current main can read standalone GGUF metadata, Qwen BPE tokenizers and embedded chat templates. Native adapters cover Qwen3.5 dense, Qwen3.5 MoE, Qwen4Exp / Flash-Next and DFlash2 drafts, with tensor-parallel storage and mixed tensor types handled explicitly: #809, #876, #888, #895.

The operator work keeps original quantization information through layout preparation, selects measured TurboMind SM70 routes and retains packaged native/reference fallbacks. Compatible IQ3_S, IQ4_XS and IQ3_XXS gate/up pairs can share activations without replacing the checkpoint with a newly quantized model.

PR #966 admits seven additional IQ3_XXS/IQ3_S mixed pairs at M8/N4352/K5120, with FP32 dot products and split-K reduction. The combined mixed-pair coverage is 18/40 layers and 48.463% of mixed-pair source bytes in that measured model. Other shapes retain canonical dispatch. The Q4_K records in #968 prepare lossless storage only; they do not add a new accelerated model route.

One retained Flash-Next GSQ-RCO IQ3_S + FP16 MTP4 development composition reports:

Workload Full round Measured decode
C1 · 8192 input / 256 output 26.072 ms 92.267 tok/s
C4 · 128 input / 1024 output per request 52.059 ms 280.441 aggregate tok/s

Evidence: #940 and the complete workload record. These use a named development build on 4× V100-SXM2-32GB / TP4, direct NVLink ring edges, FP16 activation/KV, FP32 SSM state and FULL decode graphs. Fixed-length greedy timing ignores EOS and is repeated three times; four separate natural-EOS prompts match the original outputs. Longer timing trajectories differ across versions, and the 15–17 ms round goal remains open.

Adapter support, kernel admission and model-quality qualification are separate milestones. The dense and DFlash2 adapters have recorded loading/generation checks; the MoE adapter PR alone does not establish broad 35B quality or throughput. Consult the GGUF controls and each model's recorded gates. These additions postdate the v1.5.1 release wheel.


Qwen3.6-35B-A3B NVFP4

PR #270 adds an exact SM70 route for mixed ModelOpt NVFP4 checkpoints.

Highlights:

  • FP8 dense projections;
  • W4A16_NVFP4 routed/shared experts;
  • grouped TurboMind MoE;
  • duplicate expert-slot preservation;
  • mixed-precision GDN routing;
  • MTP cold-start warmup.

Matched no-MTP:

AWQ:
prefill 0.3813 s
decode 113.71 tok/s

NVFP4:
prefill 0.4216 s
decode 116.99 tok/s

MTP4:

174.76 tok/s
1.49× NVFP4 no-MTP

Quality:

GSM8K:
122/128
95.3125%

invalid outputs:
0

repetitive records:
0

DeepSeek-V4 on V100

DeepSeek-V4 work extends beyond a single sparse-attention kernel.

The SM70 stack includes work around:

  • sparse MLA;
  • FP8 dense projections;
  • MXFP4 experts;
  • grouped MoE;
  • Indexer;
  • KPool;
  • Q normalization / RoPE / KV insertion;
  • custom TP4 all-reduce;
  • PP2×TP4 execution;
  • exact GEMV hot paths.

Representative results:

TP8 no-spec:
≈65.1 tok/s

PP2×TP4 strict quality control:
73.539 tok/s

PP2×TP4 combined endpoint:
73.613–73.646 tok/s

Strict quality control:

GSM8K    : 64/64
HumanEval: 29/32
LongBench: 44.740

GLM-5.3 on V100

The current GLM-5.3 SM70 path uses:

  • ModelOpt NVFP4 MoE;
  • FP16 non-expert weights;
  • FP8 E4M3 KV;
  • TP4 / PP2;
  • sparse MLA;
  • exact KDA GEMV;
  • fused KDA f/g;
  • mHC;
  • custom all-reduce;
  • full decode CUDA Graph.

Retained stability result:

Decode:
53.013085
53.018516
53.017527 tok/s

Mean:
53.016376 tok/s

Mean TPOT:
18.862097 ms

1K prefill:

266.039984 tok/s

The quality audit also records a reasoning-mode caveat: Max reasoning can exhaust the output budget on concise code tasks, while the targeted low-reasoning rerun completes and passes both AST and external execution checks.

The figures above remain the historical no-MTP audit. Later merged work integrates GLM DFlash2 with the current scheduler, TP8 verifier and PP2 KV-capacity handling: #501, #503. Those changes do not turn the 53.016 tok/s no-MTP record into a DFlash2 benchmark.


🎬 Native Image / Video Workflows

The repository also contains native MiniMax H3 audio/video generation and Z-Image Turbo / Base image jobs. H3 work covers SM70 attention, tensor-parallel execution, weight residency, export and actual lifecycle/denoise progress. Z-Image adds persistent asynchronous jobs, synchronous compatibility, cancellation and restart handling: #557, #583, #585.

The recorded V100 workflow checks include 1024-square Turbo/Base images and H3 text/keyframe/reference-image video at 1344×768, 107 frames and 24 fps with stereo audio. These are bounded workflow checks, not qualification of every duration, model variant or final release-wheel configuration. See native creative jobs for the tested recipes and limits.


🧠 What We Mean by “Make Volta Fast Again”

We do not claim V100 has the same theoretical peak as A100, H100, or Blackwell.

The point is different.

A large amount of modern inference software simply does not seriously optimize for SM70 anymore.

That creates two gaps:

hardware-generation gap
+
software-neglect gap

1Cat-LLM works on the second gap.

When representative Attention useful compute moves from:

17.92 TFLOP/s

to:

46–47 TFLOP/s

and then to:

≈60.8 TFLOP/s

while real 27B 256K decode still reaches:

50.376 tok/s

the conclusion is not that V100 “became A100.”

The conclusion is:

Software stopped wasting V100.


📦 Installation

Install the published v1.5.1 release into a clean environment. Download its wheel and SHA256SUMS from the same release.

Published binary requirements:

Linux x86_64 · GLIBC >= 2.38 · GLIBCXX >= 3.4.21
Python 3.12
PyTorch 2.10.0 + CUDA 12.8 runtime dependencies
CUDA 12.8 Toolkit + C++ compiler + ninja for remaining TileLang JIT
SM70 / Tesla V100

The artifact is 1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl. It has not passed the manylinux container gate. Its standard CUDA Toolkit/compiler requirements remain even though the wheel bundles the SM70 extensions, Flash-V100, FlashQLA and launchers. The release was tested with driver 580.173.02 and recommends 570.124.06 or newer.

With uv installed:

sha256sum --ignore-missing -c SHA256SUMS
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python \
  ./1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl
source .venv/bin/activate

The installer resolves declared Python/Torch dependencies. Download target and draft weights separately. The release profile needs no private developer .so, source overlay or performance environment overrides; retain the standard Toolkit/JIT setup above.

Verification:

python - <<'PY'
import sys
import torch
import vllm
import flash_attn_v100
from flash_attn_v100 import flash_attn_v100_cuda, paged_kv_utils
from flash_attn_v100 import flash_attn_grouped_verify_max_query_tokens

print("Python:", sys.version.split()[0])
print("Torch:", torch.__version__)
print("CUDA:", torch.version.cuda)
print("GPU:", torch.cuda.get_device_name(0))
print("Compute capability:", torch.cuda.get_device_capability(0))
print("vLLM:", vllm.__version__)
print("flash_attn_v100:", flash_attn_v100.__version__)
print("DFlash2 grouped verify max Q:", flash_attn_grouped_verify_max_query_tokens())
print("FlashAttention-V100: OK")
PY

This checks imports and visible hardware. Validate actual startup and requests using your model, topology and memory budget; it is not a performance test.


▶ Qwen3.8-27B-NVFP4 + DFlash2

Example TP4 + E4M3 serving command

The release wheel installs the V100 launcher and its versioned profile. Supply the target checkpoint path:

serve_qwen38_27b_nvfp4_v100.sh /models/Qwen3.8-27B-QUASAR-NVFP4

The profile pins TP4, FP16 compute/draft KV, FP8 E4M3 target KV, 256K context, --max-num-batched-tokens 8192, --max-num-seqs 4, and the 2048/8192 KV and Mamba block sizes. Append normal vllm serve options to override a release default. This profile is validated for four peer-connected V100-SXM2 32GB GPUs; other hardware and concurrency levels need a separate memory and speed check.

The first launch downloads the pinned incoai/Qwen3.8-27B-DFlash2 revision (about 3.85 GB of draft weights). To use an existing local draft:

serve_qwen38_27b_nvfp4_v100.sh /models/Qwen3.8-27B-QUASAR-NVFP4 \
  --draft /models/Qwen3.8-27B-DFlash2

For the validated Qwen3.8 DFlash2 contract, runtime policy resolves the checkpoint-native draft geometry and the SM70 draft Attention backend.

Representative automatic values:

official draft block size = 8
draft width               = 7
selector Top-K            = 16
release target KV         = FP8 E4M3
release draft KV          = FP16
draft attention backend   = FLASH_ATTN_V100
verification fast paths   = automatic

Operators select compatible paths from the actual dtype, local head/projection layout, shape and execution policy. E4M3 grouped FP32 and long-context verification are part of the release recipe. Explicit E5M2 remains available for historical reproduction, but does not select the E4M3-specific acceleration paths. A changed TP degree, KV format or service capacity needs its own memory, speed and quality check.

Inspect the installed configuration and the running server's report:

python -m vllm.sm70_profiles show qwen38_27b_nvfp4_dflash2 --json
curl --fail http://127.0.0.1:8000/v1/sm70/acceleration

Supply the server's API-key authentication if enabled. The report distinguishes configured capabilities, and current main also reports loaded kernel instances and fallback reasons (#802). A configured or prepared path is not proof that every request executed it; pair the report with worker route logs. See the release profile guide.


Historical DFlash2 release-path measurements

These retain the original pre-1.5.1 PR contracts. Use the release-to-release table above for the published v1.5.1 comparison.

Contract Result
Complete DFlash2 round ≈17.38 ms
32K cold prefill ≈4,039–4,069 tok/s
64K pure prefill ≈3,567 tok/s
32K vs retained pre-closure DFlash2 +30.1%
64K vs retained pre-closure DFlash2 +37.7%
Historical web-prompt streaming decode 206.06 tok/s
High-acceptance MBPP request 251.60 tok/s
Adaptive lookup q16 repeated-context 316.27 tok/s
Structured API 24/24 pass
Long alternating-prefix state 5/5 pass
Target-only / DFlash2 WikiText PPL 5.4993116 / 5.4993622

▶ Flash-Next NVFP4 + MTP4

The same release includes a separate Flash-Next launcher:

serve_flash_next_nvfp4_v100.sh /models/Qwen3.8-Flash-Next-NVFP4

Its defaults are TP4, FP16 compute/KV, MTP4, 131072 context, 8192-token prefill, one sequence, memory utilization 0.90, prefix caching and synchronous scheduling. It targets four peer-connected V100-SXM2-32GB GPUs and automatically selects compatible FP16 acceleration and hybrid PLE placement.

The release test machine had ample host RAM: Flash-Next PLE tables alone total approximately 47.68 GiB, before model mappings, caches and process memory. Low-RAM operation was not qualified by that release test. Later main includes capability-based device/pinned-host/disk PLE placement and packed GGUF lookup; these have their own storage and model-validation conditions (#806, #877, #885).


🔨 Build From Source

Clone:

git clone https://github.com/1CatAI/1Cat-vLLM.git
cd 1Cat-vLLM

Set the bundled CUDA extensions' SM70 build targets:

export TORCH_CUDA_ARCH_LIST=7.0
export CMAKE_CUDA_ARCHITECTURES=70

Use this fork's source checkout with CUDA 12.8 and its CUDA build dependencies. The source-build procedure describes editable builds; its upstream prebuilt-wheel examples are not the 1Cat SM70 release. Keep the bundled Flash-V100 and other native extensions matched to the source revision.

The two variables above are build-time inputs for a source build only. They are already fixed in the SM70 release-wheel build and are not needed after installing that wheel.

When building with CUDA 12.8 for SM70, dependency metadata selects the CUDA 12.8 PyTorch wheels matching the build interpreter and Linux architecture. The published SM70 release wheel still requires Python 3.12 on x86_64; building for another interpreter does not make that release wheel compatible.

Because this project contains custom CUDA extensions, make sure the active compiler/toolkit matches the PyTorch CUDA ABI used by your environment.


📐 Benchmarking Policy

1Cat-LLM intentionally separates:

kernel latency
operator throughput
Attention useful TFLOP/s
prefill tok/s
target-only pure decode
speculative pure decode
streaming decode
endpoint throughput
task-quality score
PPL / distribution checks

A benchmark claim is most useful when it retains:

  • exact model/checkpoint;
  • GPU type/count;
  • TP/PP topology;
  • context and output length;
  • batch size;
  • KV dtype;
  • quantization route;
  • CUDA Graph mode;
  • prefix-cache state;
  • sampling contract;
  • speculative method;
  • acceptance length;
  • quality result;
  • whether the result is measured or projected.

This README follows that policy wherever the underlying PR retained enough information.


🛡️ Promotion Policy

A fast path is not promoted solely because a microbenchmark is faster.

Depending on the arithmetic change, promotion may require:

  • bitwise operator equality;
  • bounded numerical error;
  • CUDA Graph replay stability;
  • same-contract endpoint speed;
  • dataset quality;
  • natural-stop / output-health checks;
  • PPL / logprob distribution checks;
  • explicit rollback;
  • structural/runtime admission rather than hard-coded model identity.

An open or Draft PR is not a merged feature. A merged PR may also retain rejected experiments alongside its accepted implementation; merge status does not qualify every number in its history.

The early ≈79 TFLOP/s Attention candidate is a good example: its historical 256K quality failure does not qualify that arithmetic as a production default. Later repaired FP32 routes were merged separately. The original 79T number remains historical; the later full-FP32 measurements keep their own precision and validation contracts.

Open Follow-ups at This Snapshot

PR Reported scope Status
#956 Superseded Mamba state-block retention with align mode, async scheduling and prefill chunks spanning multiple state blocks Open; pending integration
#779 Bound TurboQuant SDPA continuation scratch; retain numerical and model-quality gates Draft; pending integration
#715 DeepSeek-V4 FP16 hc_head range handling for attention-sink rows Open; pending integration

These reports remain separate from the merged improvements above. Check the linked PR for the affected configuration, evidence and current resolution before relying on that fix.


🧱 Runtime, Not Just Kernels

1Cat-LLM includes work across the whole serving path:

  • FlashAttention-V100;
  • paged KV utilities;
  • FP8 KV bridges;
  • QSA sparse Attention;
  • FlashQLA / GDN;
  • TurboMind SM70 quantized kernels;
  • grouped MoE;
  • MTP;
  • DFlash2;
  • CUDA Graph;
  • prefix cache;
  • hybrid Mamba state;
  • custom all-reduce;
  • sampling;
  • tool calling;
  • reasoning parser;
  • structured output;
  • standalone GGUF loading and mixed-format projection preparation;
  • capability-based kernel selection and loaded-kernel diagnostics;
  • mixed-prefill scheduling and hybrid prefix replay;
  • device / pinned-host / disk PLE storage;
  • native image / audio / video jobs;
  • wheel / RPATH / ABI packaging.

A fast kernel is only useful if the full model and serving API can use it correctly.


💾 More Room for the Model and Its Context

Recent work also reduces resident weights, scale copies, graph scratch and loading peaks. Representative PR measurements, per GPU, include:

Change Recorded result Scope
Shared weights / codes / scales Model allocation 10.07 → 6.47 GiB; available KV budget 10.57 → 12.68 GiB #671 · QUASAR NVFP4 / TP4 / DFlash2 q7 / E4M3 target KV
Shared graph scratch / draft snapshots Non-KV active memory: TP2 17.117 → 15.146 GiB, TP4 10.225 → 8.794 GiB #677 · recorded default ON/OFF comparison
Reused compiled subgraphs First startup 271.26 s, matching-config restart 95.30 s v1.5.1 · two observations within the new version

The memory records use V100-SXM2-32GB, CUDA 12.8 / Torch 2.10, FP16 compute/draft KV, 262144 context, 8192-token prefill and CUDA Graphs. #671 uses 32 sequence slots / memory .85; #677 uses 32 slots / memory .90. The savings overlap and must not be added. Automatic KV allocation can consume the recovered space, so total NVML usage need not fall by the same amount. Restart caching still loads weights and captures CUDA graphs; it does not make startup free.

Local tensor geometry now also admits qualified TP1 / TP2 paths. #666 records target-only 27B NVFP4 checks at TP1 64K and TP2 256K on 32GB cards. These are bounded capability checks, not a claim that the TP4 DFlash2 release recipe fits unchanged on one card. Higher client concurrency can queue behind the actual resident capacity.


🧭 Project Direction

1Cat-LLM focuses on a simple question:

How much modern LLM inference performance is still hidden inside Volta if the software stack is redesigned instead of abandoned?

Current directions include:

  • further long-context Attention work;
  • lower DFlash2 verifier cost;
  • higher-acceptance speculative execution;
  • sparse Attention;
  • modern quantization formats on SM70;
  • fused decode hot paths;
  • MoE routing and grouped GEMM;
  • broader GGUF model-quality and serving qualification;
  • mixed-load latency, cache reuse and memory headroom;
  • multi-model SM70 support;
  • stable wheel/release packaging.

💬 WeChat Community

Join the 1Cat-LLM Open-Source Community by adding WeChat ID YM_isi to request the latest group invitation.

1Cat-vLLM WeChat Group 8 QR code


❤️ Acknowledgements

1Cat-LLM builds on the work of the broader open-source inference ecosystem, including vLLM, NVIDIA CUDA, FlashAttention, CUTLASS/TurboMind-related kernels, model authors, quantization projects, and contributors whose work is referenced in individual PRs and source files.

Special thanks to @yangzhuxinyzx and @1CatTCat for their outstanding contributions to the continued evolution and performance breakthroughs of 1Cat-LLM.

Where external implementations or algorithms are adapted, provenance and license information should be preserved in the corresponding source and PR history.


License

Please refer to the repository license and the licenses of bundled or adapted third-party components.