Research · to

ORCA: Evaluating LLMs on Data Science Code Translation

76Developing2 outlets · 2 reportsarXiv cs.AIHF Daily Papers

The story

AI · 2 outlets

Why it mattersHelps developers accurately assess LLM cross-library code translation ability.

ORCA benchmark launched to fill evaluation gap for cross-library translation of data science code by large models

HF Daily Papers published a paper on constructing LLM alignment data from Islamic ethics

What happened

Addressing the lack of research on evaluating cross-library translation of data science code by large models, a research team built the ORCA comprehensive benchmark, containing 1,600 fine-grained tasks and 200 full project-level tasks, paired with reference translations and functional verification test cases, and quality-verified through multiple stages. Tests show that the frontier model Claude-Opus-4.6 achieved translation success rates of 56.92% and 33.67% on the two task types respectively; the intent-enhanced translation method proposed by the team delivered absolute success rate gains of 4.80% and 5.33% respectively.

Key facts

Publisher
Research team (published publicly on arXiv cs.AI on September 28, 2026)
Product
ORCA comprehensive benchmark for data science code translation
Scale/Composition
ORCA-MAIN contains 1,600 fine-grained tasks; ORCA-PROJECT contains 200 full project-level tasks
Release Date
September 28, 2026
Benchmark Results
Claude-Opus-4.6 translation success rates of 56.92% and 33.67% on the two subtask sets respectively
Method Gain
Intent enhancement method improved success rates by 4.80% and 5.33% on the two subtask sets respectively

Background

There is a notable research gap in specialized evaluation of large models for cross-library translation of data science code, which requires functional equivalence and ecosystem interoperability, with no mature unified evaluation framework currently available.

Why it matters

For the industry, ORCA fills the evaluation gap in data science code translation and can precisely identify capability shortcomings of large models in cross-ecosystem code adaptation. For developers, the benchmark can help select large model tools suited to multi-library migration needs, and the intent enhancement optimization approach can also be directly applied to daily translation tools. For general users, improved cross-library code translation capabilities can effectively lower the barrier to using multiple data science ecosystems.

What to watch

Worth watching going forward are the iterative performance of frontier large models on the ORCA benchmark and progress in the practical application of the intent enhancement method.

Written by AI from 2 reports and updated as new ones arrive. It may contain mistakes; the original is the source of truth.

Coverage timeline

Cross-checked: 2 independent outlets (arXiv, Hugging Face Papers) covered this; several channels of one company count once. The score gets a 10-point bonus on top of the best single report.

  1. arXiv cs.AI ↗ORCA: Evaluating LLMs on Data Science Code TranslationProposes ORCA benchmark to evaluate LLM performance on data science code translation.
  2. HF Daily Papers ↗Constructing Alignment Data from Normative FrameworksMethod builds LLM alignment data based on Islamic ethical frameworks.
Back to AI News