Daily timeline · Sep 28 · 北京时间 UTC+8
AI news on 2026-09-28
76 items, 35 of them scored 70+. Days follow Beijing time (UTC+8) and items are ordered by when this site picked them up, so late-publishing sources still land on the right day. Click a title for details; X posts open on X.
- 量子位70Industry清华系量子AI团队估值10亿,改造大模型底层清华背景量子AI创业团队获10亿估值,用量子技术优化大模型底层架构
- 量子位70Products3 outlets桌面级量子计算小盒子发布,端到端跑通数据不出门桌面量子计算设备已跑通端到端流程,数据无需外传,降低开发者使用门槛
- 量子位70Models匿名玉兔模型登顶OpenRouter双榜,Coding能力优异玉兔模型登顶OpenRouter双榜,中秋稳坐日榜首,编程实测表现突出Published ,picked up
- TechCrunch AI83IndustryAnthropic’s CEO is about to have dinner with President TrumpAnthropic首席执行官Dario Amodei即将与特朗普共进晚餐,此为双方首次一对一单独会面。
- The Verge AI74Industry2 outletsOpenAI agents tried to ‘bruteforce’ a UN website安全研究员发现OpenAI智能体4-6月扫描联合国贸发会议网站超1.6万次,凸显AI智能体安全风险。Published ,picked up
- The Verge AI67ProductsEngram is a sampler that turns broken AI hallucinations into music该AI采样器可将AI幻觉生成的音频转为音色,并非一键生成歌曲设备,正开启众筹。
- TechCrunch AI52IndustryCan Muse overcome Meta’s trust issues?Equity播客讨论Meta的AI公告抢占OpenAI、Anthropic风头的现象及Muse前景。
- 量子位42Industry4 outlets量子位称存在可致OpenAI最强模型训练崩溃的题目该资讯未披露具体题目及技术细节,仅提及存在可致OpenAI最强模型训练崩溃的题目。Published ,picked up
- TechCrunch AI28IndustryAnthropic’s Dario Amodei gets the SNL treatment美国综艺《周六夜现场》推出恶搞Amodei的桥段,调侃其为AI的"造魔者"。Published ,picked up
- arXiv cs.AI35Research2 outletsTowards an Implementation Architecture for the S3Q Machine Qualia TheoryProposes a five-layer computational implementation architecture for the S3Q machine consciousness theory
- arXiv cs.CL82ResearchAuditing LLM-as-Judge Failures in Production Text-to-SQL PipelinesIt finds low agreement between production LLM judges and human annotators, proposing fixes.
- arXiv cs.CL82Research2 outletsCartograph: Federated Tool Discovery Framework for AI AgentsIt reduces AI agent tool catalog traversal complexity from O(n) to O(k) via federated proxy.
- arXiv cs.CL82Research2 outletsMemProbe: Diagnosing Stability-Plasticity Tradeoffs in Agent MemoryIt proposes MemProbe, a cognitive-inspired framework to diagnose LLM agent memory tradeoffs.
- arXiv cs.CL82Research3 outletsNew Benchmark for Web Agents' Knowledge Synthesis CapabilitiesIt proposes a new benchmark to evaluate web agents' knowledge synthesis skills.
- arXiv cs.AI82ResearchLLM Parkinsonism: Executive Control Failure in Autonomous LLM AgentsIdentifies LLM agent executive control flaws, proposes uncertainty-aware control architecture.
- arXiv cs.CL80ResearchSpotify's Bootstrapping Method for Conversational Recommendation AgentsSpotify shares synthetic data and self-improvement loops for its conversational recommendation agents.
- arXiv cs.CL80ResearchFailure Analysis of Retrieval-Based Evaluation for Medical LLM AnswersIt analyzes failure modes of retrieval-based fact verification for medical LLM answers.
- arXiv cs.CL80ResearchCARGO: Context-Aware Evaluation Framework for Production AI AgentsIt proposes a context-aware evaluation framework to reduce misjudgment for production AI agents.
- arXiv cs.CL77ResearchTarget Speaker Unlearning for LLM-Based ASR at Inference TimeIt proposes TSU-ASR to skip transcription of opt-out speakers during inference.
- arXiv cs.CL77ResearchCode-Switching Curricula Improve Cross-Lingual Alignment in Small LMsIt finds code-switched text training induces cross-lingual alignment in small Transformer models.
- arXiv cs.CL77ResearchBeyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache EvictionThe paper proposes a new KV cache eviction scoring method integrating attention diversity and redundancy.
- arXiv cs.CL75Research2 outletsHierarchical Collaborative Memory for LLM Agent RetrievalIt proposes a validity-aware retrieval mechanism for LLM agents' heterogeneous memories.
- arXiv cs.CL75ResearchSlideLab: Audience-Centered Scientific Slide Generation FrameworkIt's a training-free multi-agent framework that generates scientific slides from research papers.
- arXiv cs.CL75Research2 outletsCausality-Aware LLM Framework for Simultaneous Speech TranslationIt addresses causal alignment data scarcity for LLM-based simultaneous speech translation.
- arXiv cs.CL75Research2 outletsInquesto Score: Reliability Evaluation Protocol for Voice AgentsIt proposes a reproducible, interpretable reliability evaluation protocol for voice agents.
- arXiv cs.AI75ResearchBioEVAL: Global Multi-Institution Benchmark for Bioengineering AI ModelsBioEVAL is a global multi-institution benchmark for evaluating LLMs and multimodal models in bioengineering.
- arXiv cs.AI75ResearchAudio LLMs Know When They Can't Hear YouStudies audio LLM's ability to recognize unreliable self-transcription of degraded audio.
- arXiv cs.AI75ResearchCRC-Router: Risk-Constrained Routing for Medical Agentic AI SystemsProposes risk-constrained routing to ensure safe deployment of medical AI agents.
- arXiv cs.AI75ResearchAnalyzing and Mitigating Cost-Inefficient Behaviors in Coding AgentsAnalyzes cost-inefficient behaviors of coding agents, proposes corresponding mitigation strategies.
- arXiv cs.AI75Research2 outletsLearning What to Skip for Efficient Multi-Agent LLM WorkflowsProposes counterfactual credit assignment to skip redundant steps in multi-agent LLM workflows.
- Simon Willison74Tips & viewsSimon Willison Releases 2026 LLM Trends Keynote ResourcesHe delivered a closing keynote on 2026 LLM trends, sharing accompanying video, slides and notes.Published ,picked up
- arXiv cs.CL73ResearchSurvey on Fake Review Detection: From PLMs to LLMsIt reviews fake review detection evolution and LLM's dual impact on the field.
- arXiv cs.CL73ResearchDiversifying Personas to Reduce LLM Output HomogeneityIt studies persona diversification to reduce LLM output homogeneity and groupthink.
- arXiv cs.CL73ResearchI-Parakeet: Integer-Only Conformer ASR on Mobile NPUI-Parakeet is an integer-only Conformer ASR running fully on mobile NPUs without floating-point operators.
- arXiv cs.AI73ResearchSkill Cascading Attacks on Open Skill-Based AI Agent SystemsThe paper reveals a new attack path where malicious skills on agent platforms cause cascading hidden harms.
- arXiv cs.AI73ResearchBenchy: A Universal Semantic Language for Task-Oriented AI BenchmarksBenchy is a semantic language and execution engine for standardized, portable AI benchmark definition.
- Simon Willison70ProductsSimon Willison Releases Bluesky Reply Bot Checker ToolHe launched a tool to detect automated reply bots on Bluesky using its open API.Published ,picked up
- arXiv cs.CL69ResearchMechanistic Study of BERT Neurons for AI Text DetectionThe study analyzes neurons in frozen BERT that support AI-generated text detection.
- arXiv cs.CL69Research2 outletsManifold Projection for Masked Language Modeling OptimizationIt proposes a new attention-free context mixing scheme for masked language models.
- arXiv cs.CL69ResearchBenchmark Framework for Systematic Review Screening AutomationIt proposes a benchmark for evaluating LLM-assisted systematic review article screening.
- arXiv cs.CL69ResearchREALMS: Conversational AI System for Real-Time Audience SizingIt provides real-time exact audience sizing for high-dimensional nested profiles in marketing.
- arXiv cs.AI69ResearchScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?ScopeBench is a benchmark testing whether autonomous AI agents break authorized boundaries under pressure.
- arXiv cs.AI69ResearchT-RoPE: Time-Aware Rotary Position Embedding for Sequential RecommendationProposes time-aware rotary position embedding adapted for sequential recommendation scenarios.
- arXiv cs.AI69ResearchBAER: Backbone-Adaptive Evidence Routing for LLM JudgingProposes BAER adaptive routing method to improve robustness of pairwise LLM judging.
- arXiv cs.AI68Research2 outletsHARDEN: Generating Harder Answer-Preserving Test Cases via Evolutionary SearchHARDEN is a constrained evolutionary search method to generate harder test cases preserving original outputs.
- arXiv cs.CL66ResearchSignTrace: Reverse Lookup System for Chinese Sign LanguageIt supports natural language queries for Chinese sign language meanings via LLM tools.
- arXiv cs.AI66ResearchMulti-Agent Code Judge Reliability: Label-Free Measures and Decline MechanismThe paper proposes label-free reliability measures and a code judge that avoids unfounded guesses.
- arXiv cs.AI66Research2 outletsMCP-Based Architectural Mediation for LLM Agents Connecting to Data SpacesThe paper proposes an MCP-based mediation approach to solve LLM agent integration with governed data spaces.
- arXiv cs.AI66ResearchKnowledge Graph-Based Framework for Evaluating LLM Context UnderstandingThe paper proposes a knowledge graph-based framework to quantitatively evaluate LLM's context comprehension.
- arXiv cs.AI66ResearchLAVOIR: Teaching Single-Pass Decision Encoders to Ask for InfoProposes LAVOIR mechanism for single-pass encoders to actively ask for missing information.
- arXiv cs.AI66Research2 outletsORCA: Evaluating LLMs on Data Science Code TranslationProposes ORCA benchmark to evaluate LLM performance on data science code translation.
- arXiv cs.AI66ResearchHCOE: Hyperbolic Clinical Ontology Embeddings for Biomedical LMsProposes HCOE hyperbolic embedding to preserve medical code hierarchy structure.
- arXiv cs.CL64Research2 outletsRecursive Self-Improvement via On-Policy Distillation for ReasoningThe paper proposes on-policy self-distillation to remove external teacher demand for reasoning model improvement.
- arXiv cs.AI64ResearchIntuitive Prompting Improves LLM Agent Social Media Reaction Simulation FidelityThe paper finds intuitive prompting improves LLM agent fidelity when simulating social media user reactions.
- arXiv cs.CL62ResearchSEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian LanguagesSEA-CLIP-Tiny is a compact image-text embedding model tailored for Southeast Asian low-resource languages.
- arXiv cs.CL62Research2 outletsUnderstanding the Role of Prompt Template in Knowledge Distillation for Safety AlignmentThe paper studies how prompt template choice impacts student model safety alignment robustness in distillation.
- arXiv cs.AI62ResearchThe Price of Thought: Does Test-Time Reasoning Pay in LLM TradingEvaluates test-time reasoning cost vs. return for LLM-based quantitative trading systems.
- arXiv cs.CL60ResearchTRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video UnderstandingTRACE is a new evaluation framework addressing gaps in current streaming video understanding assessment.
- arXiv cs.AI60ResearchSpectral Feedback for Test-Time Alignment of Protein Diffusion ModelsThe paper proposes a spectral feedback method to improve test-time alignment of protein diffusion models.
- arXiv cs.CL58ResearchLearning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized GuidanceThe paper proposes a randomized guidance method to improve tandem speech model conversational naturalness.
- arXiv cs.AI58ResearchSynthetic Ground-Truth Evaluation Framework for Explainable AI MethodsThe paper proposes a synthetic ground-truth framework to address evaluation gaps for explainable AI methods.
- arXiv cs.AI58Research2 outletsTeacher-Guided Fitness Approximation for Efficient TinyML Architecture SearchThe paper proposes a teacher-guided low-fidelity framework to reduce TinyML architecture search computation cost.
- arXiv cs.AI56ResearchBringing AI to Autonomous Systems -- From Cognition to Collective IntelligenceThe paper proposes a design framework for AI-powered autonomous systems combining connectionist and symbolic AI.
- arXiv cs.AI56ResearchSelective Amortization for Efficient Visual Token CommunicationProposes selective reasoning amortization to reduce cost of visual token communication.
- arXiv cs.CL55ResearchWords Speak Louder Than Order: A Behavioral Evaluation of Gemma 4The paper evaluates Gemma 4's decision preference when presented with conflicting input documents.
- arXiv cs.CL55Research2 outletsEnhancing Assessment of Self-Consistency in LLM Explanations using Perturbation StrengthThe paper proposes a unified perturbation strength measure to improve LLM explanation self-consistency assessment.
- arXiv cs.AI55ResearchTransmembrane Protein Topology Prediction from 3D Structure with GNNThe paper presents a GNN-based approach to predict transmembrane protein topology from 3D structural data.
- arXiv cs.AI55ResearchPretrained ASR Pseudo-labeling for Noisy Police Audio TranscriptionThe paper evaluates pseudo-labeling efficacy for ASR adaptation on noisy police communication audio.
- arXiv cs.AI55ResearchAtelier: Self-Supervised CryoEM Volume Feature Learning via HypernetworksThe paper proposes Atelier, a hypernetwork-based self-supervised method for cryoEM volume feature learning.
- arXiv cs.AI54ResearchInsurance Reserve Intelligence Platform for Actuarial EstimationProposes AI-powered platform to simplify insurance reserve estimation calculations.
- arXiv cs.CL51ResearchUnified Account of Concepts and Chunks in CognitionIt reviews the Cobweb model to unify disjoint research on concepts and chunks.
- MIT Tech Review83Industry4 outletsWho’s liable when AI agents go rogue?MIT Tech Review explores AI agent liability amid recent agent-driven cyberattacks.Published ,picked up
- 量子位76IndustryHuawei Redefines AIDC with Computing-Electricity SynergyHuawei proposes computing-electricity synergy for next-generation AIDC AI infrastructure.
- Hugging Face75ModelsHolo4: powering generalist computer-use agentsHugging Face releases Holo4 to support generalist computer-use AI agent development.Published ,picked up
- 量子位64IndustrySiemens Xcelerator Ecosystem Empowerment BreakdownQuantum Bit breaks down Siemens' support for partners building AI Agents and going global.
- Simon Willison33Tips & viewsMuse AI Agent Auto-Reply Causes Pickup MishapSimon Willison shares how his Muse AI auto-reply led to a pickup no-show and bad rating.Published ,picked up