Daily timeline · Sep 28 · 北京时间 UTC+8

AI news on 2026-09-28

76 items, 35 of them scored 70+. Days follow Beijing time (UTC+8) and items are ordered by when this site picked them up, so late-publishing sources still land on the right day. Click a title for details; X posts open on X.

  1. 量子位70Industry清华系量子AI团队估值10亿,改造大模型底层清华背景量子AI创业团队获10亿估值,用量子技术优化大模型底层架构
  2. 量子位70Products3 outlets桌面级量子计算小盒子发布,端到端跑通数据不出门桌面量子计算设备已跑通端到端流程,数据无需外传,降低开发者使用门槛
  3. 量子位70Models匿名玉兔模型登顶OpenRouter双榜,Coding能力优异玉兔模型登顶OpenRouter双榜,中秋稳坐日榜首,编程实测表现突出Published ,picked up
  4. TechCrunch AI83IndustryAnthropic’s CEO is about to have dinner with President TrumpAnthropic首席执行官Dario Amodei即将与特朗普共进晚餐,此为双方首次一对一单独会面。
  5. The Verge AI74Industry2 outletsOpenAI agents tried to ‘bruteforce’ a UN website安全研究员发现OpenAI智能体4-6月扫描联合国贸发会议网站超1.6万次,凸显AI智能体安全风险。Published ,picked up
  6. The Verge AI67ProductsEngram is a sampler that turns broken AI hallucinations into music该AI采样器可将AI幻觉生成的音频转为音色,并非一键生成歌曲设备,正开启众筹。
  7. TechCrunch AI52IndustryCan Muse overcome Meta’s trust issues?Equity播客讨论Meta的AI公告抢占OpenAI、Anthropic风头的现象及Muse前景。
  8. 量子位42Industry4 outlets量子位称存在可致OpenAI最强模型训练崩溃的题目该资讯未披露具体题目及技术细节,仅提及存在可致OpenAI最强模型训练崩溃的题目。Published ,picked up
  9. TechCrunch AI28IndustryAnthropic’s Dario Amodei gets the SNL treatment美国综艺《周六夜现场》推出恶搞Amodei的桥段,调侃其为AI的"造魔者"。Published ,picked up
  10. arXiv cs.AI35Research2 outletsTowards an Implementation Architecture for the S3Q Machine Qualia TheoryProposes a five-layer computational implementation architecture for the S3Q machine consciousness theory
  11. arXiv cs.CL82ResearchAuditing LLM-as-Judge Failures in Production Text-to-SQL PipelinesIt finds low agreement between production LLM judges and human annotators, proposing fixes.
  12. arXiv cs.CL82Research2 outletsCartograph: Federated Tool Discovery Framework for AI AgentsIt reduces AI agent tool catalog traversal complexity from O(n) to O(k) via federated proxy.
  13. arXiv cs.CL82Research2 outletsMemProbe: Diagnosing Stability-Plasticity Tradeoffs in Agent MemoryIt proposes MemProbe, a cognitive-inspired framework to diagnose LLM agent memory tradeoffs.
  14. arXiv cs.CL82Research3 outletsNew Benchmark for Web Agents' Knowledge Synthesis CapabilitiesIt proposes a new benchmark to evaluate web agents' knowledge synthesis skills.
  15. arXiv cs.AI82ResearchLLM Parkinsonism: Executive Control Failure in Autonomous LLM AgentsIdentifies LLM agent executive control flaws, proposes uncertainty-aware control architecture.
  16. arXiv cs.CL80ResearchSpotify's Bootstrapping Method for Conversational Recommendation AgentsSpotify shares synthetic data and self-improvement loops for its conversational recommendation agents.
  17. arXiv cs.CL80ResearchFailure Analysis of Retrieval-Based Evaluation for Medical LLM AnswersIt analyzes failure modes of retrieval-based fact verification for medical LLM answers.
  18. arXiv cs.CL80ResearchCARGO: Context-Aware Evaluation Framework for Production AI AgentsIt proposes a context-aware evaluation framework to reduce misjudgment for production AI agents.
  19. arXiv cs.CL77ResearchTarget Speaker Unlearning for LLM-Based ASR at Inference TimeIt proposes TSU-ASR to skip transcription of opt-out speakers during inference.
  20. arXiv cs.CL77ResearchCode-Switching Curricula Improve Cross-Lingual Alignment in Small LMsIt finds code-switched text training induces cross-lingual alignment in small Transformer models.
  21. arXiv cs.CL77ResearchBeyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache EvictionThe paper proposes a new KV cache eviction scoring method integrating attention diversity and redundancy.
  22. arXiv cs.CL75Research2 outletsHierarchical Collaborative Memory for LLM Agent RetrievalIt proposes a validity-aware retrieval mechanism for LLM agents' heterogeneous memories.
  23. arXiv cs.CL75ResearchSlideLab: Audience-Centered Scientific Slide Generation FrameworkIt's a training-free multi-agent framework that generates scientific slides from research papers.
  24. arXiv cs.CL75Research2 outletsCausality-Aware LLM Framework for Simultaneous Speech TranslationIt addresses causal alignment data scarcity for LLM-based simultaneous speech translation.
  25. arXiv cs.CL75Research2 outletsInquesto Score: Reliability Evaluation Protocol for Voice AgentsIt proposes a reproducible, interpretable reliability evaluation protocol for voice agents.
  26. arXiv cs.AI75ResearchBioEVAL: Global Multi-Institution Benchmark for Bioengineering AI ModelsBioEVAL is a global multi-institution benchmark for evaluating LLMs and multimodal models in bioengineering.
  27. arXiv cs.AI75ResearchAudio LLMs Know When They Can't Hear YouStudies audio LLM's ability to recognize unreliable self-transcription of degraded audio.
  28. arXiv cs.AI75ResearchCRC-Router: Risk-Constrained Routing for Medical Agentic AI SystemsProposes risk-constrained routing to ensure safe deployment of medical AI agents.
  29. arXiv cs.AI75ResearchAnalyzing and Mitigating Cost-Inefficient Behaviors in Coding AgentsAnalyzes cost-inefficient behaviors of coding agents, proposes corresponding mitigation strategies.
  30. arXiv cs.AI75Research2 outletsLearning What to Skip for Efficient Multi-Agent LLM WorkflowsProposes counterfactual credit assignment to skip redundant steps in multi-agent LLM workflows.
  31. Simon Willison74Tips & viewsSimon Willison Releases 2026 LLM Trends Keynote ResourcesHe delivered a closing keynote on 2026 LLM trends, sharing accompanying video, slides and notes.Published ,picked up
  32. arXiv cs.CL73ResearchSurvey on Fake Review Detection: From PLMs to LLMsIt reviews fake review detection evolution and LLM's dual impact on the field.
  33. arXiv cs.CL73ResearchDiversifying Personas to Reduce LLM Output HomogeneityIt studies persona diversification to reduce LLM output homogeneity and groupthink.
  34. arXiv cs.CL73ResearchI-Parakeet: Integer-Only Conformer ASR on Mobile NPUI-Parakeet is an integer-only Conformer ASR running fully on mobile NPUs without floating-point operators.
  35. arXiv cs.AI73ResearchSkill Cascading Attacks on Open Skill-Based AI Agent SystemsThe paper reveals a new attack path where malicious skills on agent platforms cause cascading hidden harms.
  36. arXiv cs.AI73ResearchBenchy: A Universal Semantic Language for Task-Oriented AI BenchmarksBenchy is a semantic language and execution engine for standardized, portable AI benchmark definition.
  37. Simon Willison70ProductsSimon Willison Releases Bluesky Reply Bot Checker ToolHe launched a tool to detect automated reply bots on Bluesky using its open API.Published ,picked up
  38. arXiv cs.CL69ResearchMechanistic Study of BERT Neurons for AI Text DetectionThe study analyzes neurons in frozen BERT that support AI-generated text detection.
  39. arXiv cs.CL69Research2 outletsManifold Projection for Masked Language Modeling OptimizationIt proposes a new attention-free context mixing scheme for masked language models.
  40. arXiv cs.CL69ResearchBenchmark Framework for Systematic Review Screening AutomationIt proposes a benchmark for evaluating LLM-assisted systematic review article screening.
  41. arXiv cs.CL69ResearchREALMS: Conversational AI System for Real-Time Audience SizingIt provides real-time exact audience sizing for high-dimensional nested profiles in marketing.
  42. arXiv cs.AI69ResearchScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?ScopeBench is a benchmark testing whether autonomous AI agents break authorized boundaries under pressure.
  43. arXiv cs.AI69ResearchT-RoPE: Time-Aware Rotary Position Embedding for Sequential RecommendationProposes time-aware rotary position embedding adapted for sequential recommendation scenarios.
  44. arXiv cs.AI69ResearchBAER: Backbone-Adaptive Evidence Routing for LLM JudgingProposes BAER adaptive routing method to improve robustness of pairwise LLM judging.
  45. arXiv cs.AI68Research2 outletsHARDEN: Generating Harder Answer-Preserving Test Cases via Evolutionary SearchHARDEN is a constrained evolutionary search method to generate harder test cases preserving original outputs.
  46. arXiv cs.CL66ResearchSignTrace: Reverse Lookup System for Chinese Sign LanguageIt supports natural language queries for Chinese sign language meanings via LLM tools.
  47. arXiv cs.AI66ResearchMulti-Agent Code Judge Reliability: Label-Free Measures and Decline MechanismThe paper proposes label-free reliability measures and a code judge that avoids unfounded guesses.
  48. arXiv cs.AI66Research2 outletsMCP-Based Architectural Mediation for LLM Agents Connecting to Data SpacesThe paper proposes an MCP-based mediation approach to solve LLM agent integration with governed data spaces.
  49. arXiv cs.AI66ResearchKnowledge Graph-Based Framework for Evaluating LLM Context UnderstandingThe paper proposes a knowledge graph-based framework to quantitatively evaluate LLM's context comprehension.
  50. arXiv cs.AI66ResearchLAVOIR: Teaching Single-Pass Decision Encoders to Ask for InfoProposes LAVOIR mechanism for single-pass encoders to actively ask for missing information.
  51. arXiv cs.AI66Research2 outletsORCA: Evaluating LLMs on Data Science Code TranslationProposes ORCA benchmark to evaluate LLM performance on data science code translation.
  52. arXiv cs.AI66ResearchHCOE: Hyperbolic Clinical Ontology Embeddings for Biomedical LMsProposes HCOE hyperbolic embedding to preserve medical code hierarchy structure.
  53. arXiv cs.CL64Research2 outletsRecursive Self-Improvement via On-Policy Distillation for ReasoningThe paper proposes on-policy self-distillation to remove external teacher demand for reasoning model improvement.
  54. arXiv cs.AI64ResearchIntuitive Prompting Improves LLM Agent Social Media Reaction Simulation FidelityThe paper finds intuitive prompting improves LLM agent fidelity when simulating social media user reactions.
  55. arXiv cs.CL62ResearchSEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian LanguagesSEA-CLIP-Tiny is a compact image-text embedding model tailored for Southeast Asian low-resource languages.
  56. arXiv cs.CL62Research2 outletsUnderstanding the Role of Prompt Template in Knowledge Distillation for Safety AlignmentThe paper studies how prompt template choice impacts student model safety alignment robustness in distillation.
  57. arXiv cs.AI62ResearchThe Price of Thought: Does Test-Time Reasoning Pay in LLM TradingEvaluates test-time reasoning cost vs. return for LLM-based quantitative trading systems.
  58. arXiv cs.CL60ResearchTRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video UnderstandingTRACE is a new evaluation framework addressing gaps in current streaming video understanding assessment.
  59. arXiv cs.AI60ResearchSpectral Feedback for Test-Time Alignment of Protein Diffusion ModelsThe paper proposes a spectral feedback method to improve test-time alignment of protein diffusion models.
  60. arXiv cs.CL58ResearchLearning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized GuidanceThe paper proposes a randomized guidance method to improve tandem speech model conversational naturalness.
  61. arXiv cs.AI58ResearchSynthetic Ground-Truth Evaluation Framework for Explainable AI MethodsThe paper proposes a synthetic ground-truth framework to address evaluation gaps for explainable AI methods.
  62. arXiv cs.AI58Research2 outletsTeacher-Guided Fitness Approximation for Efficient TinyML Architecture SearchThe paper proposes a teacher-guided low-fidelity framework to reduce TinyML architecture search computation cost.
  63. arXiv cs.AI56ResearchBringing AI to Autonomous Systems -- From Cognition to Collective IntelligenceThe paper proposes a design framework for AI-powered autonomous systems combining connectionist and symbolic AI.
  64. arXiv cs.AI56ResearchSelective Amortization for Efficient Visual Token CommunicationProposes selective reasoning amortization to reduce cost of visual token communication.
  65. arXiv cs.CL55ResearchWords Speak Louder Than Order: A Behavioral Evaluation of Gemma 4The paper evaluates Gemma 4's decision preference when presented with conflicting input documents.
  66. arXiv cs.CL55Research2 outletsEnhancing Assessment of Self-Consistency in LLM Explanations using Perturbation StrengthThe paper proposes a unified perturbation strength measure to improve LLM explanation self-consistency assessment.
  67. arXiv cs.AI55ResearchTransmembrane Protein Topology Prediction from 3D Structure with GNNThe paper presents a GNN-based approach to predict transmembrane protein topology from 3D structural data.
  68. arXiv cs.AI55ResearchPretrained ASR Pseudo-labeling for Noisy Police Audio TranscriptionThe paper evaluates pseudo-labeling efficacy for ASR adaptation on noisy police communication audio.
  69. arXiv cs.AI55ResearchAtelier: Self-Supervised CryoEM Volume Feature Learning via HypernetworksThe paper proposes Atelier, a hypernetwork-based self-supervised method for cryoEM volume feature learning.
  70. arXiv cs.AI54ResearchInsurance Reserve Intelligence Platform for Actuarial EstimationProposes AI-powered platform to simplify insurance reserve estimation calculations.
  71. arXiv cs.CL51ResearchUnified Account of Concepts and Chunks in CognitionIt reviews the Cobweb model to unify disjoint research on concepts and chunks.
  72. MIT Tech Review83Industry4 outletsWho’s liable when AI agents go rogue?MIT Tech Review explores AI agent liability amid recent agent-driven cyberattacks.Published ,picked up
  73. 量子位76IndustryHuawei Redefines AIDC with Computing-Electricity SynergyHuawei proposes computing-electricity synergy for next-generation AIDC AI infrastructure.
  74. Hugging Face75ModelsHolo4: powering generalist computer-use agentsHugging Face releases Holo4 to support generalist computer-use AI agent development.Published ,picked up
  75. 量子位64IndustrySiemens Xcelerator Ecosystem Empowerment BreakdownQuantum Bit breaks down Siemens' support for partners building AI Agents and going global.
  76. Simon Willison33Tips & viewsMuse AI Agent Auto-Reply Causes Pickup MishapSimon Willison shares how his Muse AI auto-reply led to a pickup no-show and bad rating.Published ,picked up
Back to AI News