MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 13 days ago • 96
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery Paper • 2607.06374 • Published Jul 7 • 10
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models Paper • 2607.07673 • Published Jul 8 • 14
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation Paper • 2607.09362 • Published Jul 10 • 12
NeuroCogMap Reveals Cognitive Organization of Large Language Models Paper • 2607.00397 • Published Jul 1 • 12
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published Jul 10 • 9
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Paper • 2607.11849 • Published Jul 13 • 33
Vera: A Layered Diffusion Model for Content-Preserving Video Editing Paper • 2606.23610 • Published Jun 22 • 12
BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation Paper • 2606.19651 • Published Jun 17 • 10
Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning Paper • 2606.20002 • Published Jun 18 • 10
Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills Paper • 2606.11897 • Published Jun 10 • 12
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks Paper • 2606.12871 • Published Jun 11 • 14