Research ·
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading
AI brief
AI-writtenWhy it mattersHelps quant teams assess ROI of investing in LLM inference compute.
Study shows increasing LLM inference time does not consistently improve net investment returns from stock trading
What happened
Researchers conducted controlled experiments on representative LLMs from the DeepSeek, GPT, and Gemini families, adjusting model inference investment intensity while strictly fixing information inputs, prompts, and portfolio construction rules. The tests covered a full year of U.S. stock market data, including over 800,000 asset predictions and multiple rounds of repeated model generation. Results showed that none of the models could consistently improve portfolio net returns (after deducting trading costs) by increasing inference investment.
Key facts
- Tested models
- Representative LLMs from the DeepSeek, GPT, and Gemini families
- Test scope
- Full-year U.S. stock market, covering three types of input conditions
- Test scale
- Over 800,000 asset predictions, with multiple rounds of repeated model generation
- Core test conclusion
- Increasing inference investment cannot consistently improve portfolio net returns after deducting trading costs
- Special observation
- The relationship between inference investment and trading performance for the DeepSeek model is non-monotonic
Background
The industry currently widely holds that higher reasoning investment during the LLM inference stage leads to better model decision quality, but past evaluations mostly focused on model output accuracy, rarely accounting for actual economic returns when including trading costs. Research on this front has long been lacking.
Why it matters
For the applied AI in finance industry, this conclusion breaks the entrenched belief that "stronger reasoning leads to higher returns", preventing the industry from blindly paying for high inference costs. For developers, cost-benefit checks specific to individual tasks are required before deploying LLM-based financial applications. For retail investors, the practical investment reference value of these AI tools will be subject to more pragmatic validation moving forward.
What to watch
Future work can focus on validating the actual value of LLM inference investment across different vertical scenarios, to develop scenario-specific standards for matching inference costs and benefits.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.