Research ·

DPS: Dual-Mode Precision LLM Serving System

75Developing1 reportHF Daily Papers
DPS双模式LLM服务方案提升内存利用率
Image: HF Daily Papers

AI brief

AI-written

Why it mattersCuts LLM inference service hardware costs for deployers.

The DPS system turns LLM weight memory into an elastic resource, greatly boosting serving throughput without losing accuracy.

What happened

Addressing the design limitation of existing LLM serving systems that fix model weight memory and only dynamically manage the KV cache, the research team introduced the DPS mixed-precision serving system based on a semi-unified memory mechanism. Under normal load, it runs the full-precision FP16 model. When KV cache load spikes cause GPU memory shortages, it automatically switches to a nested lower-precision weight variant, repurposing previously idle weight GPU memory for KV cache blocks. Built on vLLM and tested on dense models, Mixture-of-Experts models, and production traffic traces, the system improves sustained throughput by 2.1–3.3x compared with static FP16 deployments, boosts effective pass@1 by up to 41 percentage points, and maintains FP16-level accuracy.

Key facts

System name
DPS
Underlying platform
vLLM
Throughput improvement
2.1–3.3x
Maximum effective pass@1 improvement
41 percentage points
Core technology
Semi-unified memory with mixed-precision switching

Background

Mainstream LLM serving systems commonly use virtualization and memory optimization for KV caches, but always treat model weight memory as a fixed resource during runtime. Faced with sudden traffic spikes, they often suffer throughput drops and fail to meet service level agreements due to insufficient KV cache space.

Why it matters

For inference service providers, this solution can significantly withstand traffic spikes and reduce per-request inference costs without additional GPU memory expansion. For developers, it is compatible with the existing vLLM ecosystem and paged KV cache management mechanisms, resulting in low migration costs. For general users, response latency for large model services during peak periods will drop noticeably.

What to watch

Follow-up work will observe whether this mixed-precision elastic memory approach is officially merged into mainstream inference frameworks such as vLLM, and further verify long-term stability in production clusters.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. HF Daily Papers ↗DPS: Dual-Mode Precision LLM Serving SystemNew LLM serving system optimizes memory utilization under bursty loads.
Back to AI News