How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus Paper • 2609.15504 • Published 1 day ago • 20 • 3
Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents Paper • 2609.05824 • Published 11 days ago • 5 • 3
Negative Self-Distillation: Learning to Reason by Avoiding Flaws Paper • 2609.11699 • Published 6 days ago • 35 • 3
Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2 Paper • 2608.15380 • Published 10 days ago • 23 • 3
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Paper • 2609.10016 • Published 7 days ago • 33 • 5
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 8 days ago • 28 • 4
Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model Paper • 2609.06008 • Published 11 days ago • 21 • 4
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 13 days ago • 149 • 13
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization Paper • 2609.05258 • Published 12 days ago • 21 • 3
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space Paper • 2608.29188 • Published 18 days ago • 11 • 3
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 13 days ago • 26 • 5
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 13 days ago • 383 • 4
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents Paper • 2608.30322 • Published 16 days ago • 4 • 4
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published 15 days ago • 36 • 3
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 18 days ago • 29 • 4
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 19 days ago • 106 • 5
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals Paper • 2608.19269 • Published 18 days ago • 6 • 3
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 21 days ago • 196 • 5
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 20 days ago • 35 • 4
TorchMorph: CUDA-accelerated Morphological Transforms Paper • 2608.24738 • Published 22 days ago • 4 • 3