| 20260512_meme_multi_entity_evolving_memory_evaluation | paper_analysis | memory_mechanism, long_term_memory, agent_architecture, llm, reasoning, evaluation | kaist_ai, tuebingen_ai_center, naver_ai_lab | 100 | | | MEME: Multi-entity & Evolving Memory Evaluation | memory_os_of_ai_agent, memgpt_towards_llms_as_operating_systems, disentangling_memory_reasoning_llm | [object Object] | [object Object] | md | |
| 20260520_skillgenbench_benchmarking_skill_generation_llm_agents | paper_analysis | agent_architecture, evaluation, skill_curation, llm, reasoning | quanta_alpha, sjtu, pku, nus, xjtu, tsinghua_university, sufe, ntu, ucas | 100 | | | SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents | ontology_skill_analysis | [object Object] | [object Object] | md | |
| 20260529_agent_lifespan_engineering_for_deployed_systems | paper_analysis | memory_mechanism, long_term_memory, agent_architecture, evaluation, self_evolving_agents | ut_austin | 100 | | | Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems | Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems | [object Object] | [object Object] | md | |
| 20260529_stop_comparing_llm_agents_without_disclosing_harness | paper_analysis | agent_architecture, evaluation, llm, multi_agent_systems, reasoning | tulane_university, rutgers_university, virginia_tech | 100 | | | Stop Comparing LLM Agents Without Disclosing the Harness | Stop Comparing LLM Agents Without Disclosing the Harness | [object Object] | [object Object] | md | |
| 20260601_locally_coherent_globally_incoherent_multi_component_llm_agents | paper_analysis | multi_agent_systems, reasoning, agent_architecture, evaluation | princeton_university, paleka | 100 | | | Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents | scaling_large_language_model_multi_agent_collaboration, gasim_graph_accelerated_hybrid_social_simulation, harnessing_agentic_evolution, scaling_harness_agentic_ai | [object Object] | [object Object] | md | |
| 20260603_agentcl_continual_learning_language_agents | paper_analysis | continual_learning, memory_mechanism, agent_architecture, evaluation, self_evolving_agents | agentcl, ohio_state_university | 100 | | | AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents | learning_fast_slow_llms_adapt_continually, meme_multi_entity_evolving_memory_evaluation, muse_autoskill_self_evolving_skill_memory | [object Object] | [object Object] | md | |
| 20260604_agent_memory_characterization_system_implications | paper_analysis | memory_mechanism, agent_architecture, long_term_memory, llm, evaluation | stanford_university, ku_leuven, memory_agent_bench, memgpt, graphrag | 95 | | | Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads | memgpt_towards_llms_as_operating_systems, memlineage_lineage_guided_llm_agent_memory, calmem_dual_memory_conversational_ai, unlocking_working_memory_latent_reasoning, aip_graph_representation_for_learning_and_governing_agent_skills | [object Object] | [object Object] | md | |
| 20260622_how_transparent_is_diffusiongemma | paper_analysis | llm, reasoning, agent_security, evaluation, interpretability | google_brain | 100 | | | How Transparent is DiffusionGemma? | detecting_hallucinations_semantic_entropy | [object Object] | [object Object] | md | |
| 20260624_evoarena_tracking_memory_evolution | paper_analysis | memory_mechanism, agent_architecture, llm, evaluation, embodied_ai | nus, ntu, terminalbench_2 | 100 | | | EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments | Self-Harness: Harnesses That Improve Themselves | [object Object] | [object Object] | md | |
| 20260624_world_models_in_pieces_structural_certification | paper_analysis | reinforce_learning, reasoning, agent_architecture, evaluation, foundation_agents | cuhk_shenzhen | 100 | | | World Models in Pieces: Structural Certification for General Agents | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_characterization_system_implications | [object Object] | [object Object] | md | |
| 20260705_llm_agents_social_structure_latent_objective_emergence | paper_analysis | multi_agent_systems, social_simulation, agent_architecture, reasoning, llm, evaluation | cmu, dual_channel_debate, llm_agora | 100 | | | What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates | self_harness, contagion_networks_evaluator_bias_multi_agent_llm | [object Object] | [object Object] | md | |
| 20260707_evopolicygym_evaluating_autonomous_policy_evolution | paper_analysis | embodied_ai, reinforce_learning, self_evolving_agents, evaluation, agent_architecture | cuhk_shenzhen, sjtu, tsinghua_university, ustc | 100 | | | EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments | demopsd_disagreement_modulated_policy_self_distillation, llm_agents_social_structure_latent_objective_emergence, recontext_recursive_evidence_replay, evolvenav_proactive_preflection_self_evolving_memory | [object Object] | [object Object] | md | |
| 20260708_adacurl_adaptive_curriculum_rl | paper_analysis | reinforce_learning, llm, reasoning, agent_architecture, evaluation | amap | 100 | | | AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting | | [object Object] | [object Object] | md | |
| 20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domains | paper_analysis | reinforce_learning, reward_modeling, llm, reasoning, evaluation | scale_ai | 100 | | | Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains | | [object Object] | [object Object] | md | |