Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 4 days ago • 26
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 10 days ago • 136
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 9 days ago • 89
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published 20 days ago • 90
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published 13 days ago • 44
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 14 days ago • 226
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published 13 days ago • 85
Video-Oasis: Rethinking Evaluation of Video Understanding Paper • 2603.29616 • Published 26 days ago • 65
Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 25 days ago • 141
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation Paper • 2607.06624 • Published 21 days ago • 8
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies Paper • 2607.04434 • Published 21 days ago • 15
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published 20 days ago • 56
Rank-Then-Act: Reward-Free Control from Frame-Order Progress Paper • 2607.01897 • Published 26 days ago • 7
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Paper • 2607.03451 • Published 25 days ago • 33
AlayaWorld: Long-Horizon and Playable Video World Generation Paper • 2607.06291 • Published 21 days ago • 90