Research ·

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading

62Developing1 reportarXiv cs.AI

AI brief

AI-written

Why it mattersHelps quant teams assess ROI of investing in LLM inference compute.

Study shows increasing LLM inference time does not consistently improve net investment returns from stock trading

What happened

Researchers conducted controlled experiments on representative LLMs from the DeepSeek, GPT, and Gemini families, adjusting model inference investment intensity while strictly fixing information inputs, prompts, and portfolio construction rules. The tests covered a full year of U.S. stock market data, including over 800,000 asset predictions and multiple rounds of repeated model generation. Results showed that none of the models could consistently improve portfolio net returns (after deducting trading costs) by increasing inference investment.

Key facts

Tested models
Representative LLMs from the DeepSeek, GPT, and Gemini families
Test scope
Full-year U.S. stock market, covering three types of input conditions
Test scale
Over 800,000 asset predictions, with multiple rounds of repeated model generation
Core test conclusion
Increasing inference investment cannot consistently improve portfolio net returns after deducting trading costs
Special observation
The relationship between inference investment and trading performance for the DeepSeek model is non-monotonic

Background

The industry currently widely holds that higher reasoning investment during the LLM inference stage leads to better model decision quality, but past evaluations mostly focused on model output accuracy, rarely accounting for actual economic returns when including trading costs. Research on this front has long been lacking.

Why it matters

For the applied AI in finance industry, this conclusion breaks the entrenched belief that "stronger reasoning leads to higher returns", preventing the industry from blindly paying for high inference costs. For developers, cost-benefit checks specific to individual tasks are required before deploying LLM-based financial applications. For retail investors, the practical investment reference value of these AI tools will be subject to more pragmatic validation moving forward.

What to watch

Future work can focus on validating the actual value of LLM inference investment across different vertical scenarios, to develop scenario-specific standards for matching inference costs and benefits.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. arXiv cs.AI ↗The Price of Thought: Does Test-Time Reasoning Pay in LLM TradingEvaluates test-time reasoning cost vs. return for LLM-based quantitative trading systems.
Back to AI News