Research ·
DPS: Dual-Mode Precision LLM Serving System
AI brief
AI-writtenWhy it mattersCuts LLM inference service hardware costs for deployers.
The DPS system turns LLM weight memory into an elastic resource, greatly boosting serving throughput without losing accuracy.
What happened
Addressing the design limitation of existing LLM serving systems that fix model weight memory and only dynamically manage the KV cache, the research team introduced the DPS mixed-precision serving system based on a semi-unified memory mechanism. Under normal load, it runs the full-precision FP16 model. When KV cache load spikes cause GPU memory shortages, it automatically switches to a nested lower-precision weight variant, repurposing previously idle weight GPU memory for KV cache blocks. Built on vLLM and tested on dense models, Mixture-of-Experts models, and production traffic traces, the system improves sustained throughput by 2.1–3.3x compared with static FP16 deployments, boosts effective pass@1 by up to 41 percentage points, and maintains FP16-level accuracy.
Key facts
- System name
- DPS
- Underlying platform
- vLLM
- Throughput improvement
- 2.1–3.3x
- Maximum effective pass@1 improvement
- 41 percentage points
- Core technology
- Semi-unified memory with mixed-precision switching
Background
Mainstream LLM serving systems commonly use virtualization and memory optimization for KV caches, but always treat model weight memory as a fixed resource during runtime. Faced with sudden traffic spikes, they often suffer throughput drops and fail to meet service level agreements due to insufficient KV cache space.
Why it matters
For inference service providers, this solution can significantly withstand traffic spikes and reduce per-request inference costs without additional GPU memory expansion. For developers, it is compatible with the existing vLLM ecosystem and paged KV cache management mechanisms, resulting in low migration costs. For general users, response latency for large model services during peak periods will drop noticeably.
What to watch
Follow-up work will observe whether this mixed-precision elastic memory approach is officially merged into mainstream inference frameworks such as vLLM, and further verify long-term stability in production clusters.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.