| 20260423_agentevolver_self_evolving_agent | paper_analysis | agent_architecture, reinforce_learning, llm, reasoning | tongyi_lab, appworld, bfcl_v3 | 100 | | | AgentEvolver: Towards Efficient Self-Evolving Agent System | [object Object] | [object Object] | md | | |
| 20260427_agentsquare_automatic_llm_agent_search | paper_analysis | agent_architecture, llm, reasoning, memory_mechanism, reinforce_learning | tsinghua_fib_lab, neural_architecture_search | 100 | | | AgentSquare: Automatic LLM Agent Search in Modular Design Space | [object Object] | [object Object] | md | | |
| 20260427_detecting_hallucinations_semantic_entropy | paper_analysis | llm, reasoning, agent_architecture, memory_mechanism, reinforce_learning | oatml_oxford, neural_architecture_search | 100 | | | Detecting hallucinations in large language models using semantic entropy | [object Object] | [object Object] | md | | |
| 20260428_navigating_to_objects_in_the_real_world | paper_analysis | embodied_ai, reinforce_learning, sim_to_real, semantic_navigation | meta_ai, habitat_simulator | 100 | | | Navigating to objects in the real world | [object Object] | [object Object] | md | | |
| 20260428_rubrics_as_rewards | paper_analysis | reinforce_learning, llm, reasoning, reward_modeling | scale_ai | 95 | | | Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains | [object Object] | [object Object] | md | | |
| 20260430_hallucination_survey_llm | paper_analysis | hallucination, llm, survey, rag, reinforce_learning | harbin_institute_of_technology | 100 | | | A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions | [object Object] | [object Object] | md | | catastrophic_forgetting_implicit_inference |
| 20260506_openclaw_rl_train_any_agent_simply_by_talking | paper_analysis | reinforce_learning, agent_architecture, llm, self_evolving_agents, reasoning | gen_verse | 100 | | | OpenClaw-RL: Train Any Agent Simply by Talking | [object Object] | [object Object] | md | | |
| 20260509_skillos_learning_skill_curation_self_evolving_agents | paper_analysis | agent_architecture, memory_mechanism, reinforce_learning, self_evolving_agents, skill_curation | google_cloud_ai_research, grpo, skillos, alfworld | 100 | | | SkillOS: Learning Skill Curation for Self-Evolving Agents | [object Object] | [object Object] | md | | self_evolving_agents_survey_asi, memory_os_of_ai_agent, agentevolver_self_evolving_agent, reasoningbank_scaling_agent_self_evolving_reasoning_memory |
| 20260511_rl_long_horizon_reasoning_llm_expressiveness | paper_analysis | reinforce_learning, reasoning, llm, test_time_scaling, symbolic_reasoning | purdue_university, georgia_tech, scalelogic | 100 | | | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key | [object Object] | [object Object] | md | | chain_of_thought_prompting_elicits_reasoning, reasoningbank_scaling_agent_self_evolving_reasoning_memory |
| 20260512_learning_fast_slow_llms_adapt_continually | paper_analysis | catastrophic_forgetting, continual_learning, memory_mechanism, reinforce_learning, llm | uc_berkeley | 100 | | | Learning, Fast and Slow: Towards LLMs That Adapt Continually | [object Object] | [object Object] | md | | 20260504_reasoningbank_scaling_agent_self_evolving_reasoning_memory, 20260511_rl_long_horizon_reasoning_llm_expressiveness |
| 20260515_harnessing_agentic_evolution | paper_analysis | agent_architecture, self_evolving_agents, reasoning, memory_mechanism, reinforce_learning, workflow_optimization | deepwisdom, hkust_gz, sjtu, tsinghua_university, nanyang_technological_university | 100 | | | harnessing_agentic_evolution | [object Object] | [object Object] | md | | self_evolving_agents_survey_asi, skillos_learning_skill_curation_self_evolving_agents, aflow_automating_agentic_workflow_generation |
| 20260520_amr_sd_token_level_credit_assignment | paper_analysis | reinforce_learning, reasoning, llm, reward_modeling, self_evolving_agents | grpo, dapo, meituan, sciknoweval | 100 | | | AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment | [object Object] | [object Object] | md | | rl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents |
| 20260524_vector_policy_optimization_vpo | paper_analysis | reinforce_learning, llm, test_time_scaling, reasoning, reward_modeling | mit, grpo | 100 | | | Vector Policy Optimization: Training for Diversity Improves Test-Time Search | [object Object] | [object Object] | md | | |
| 20260525_gated_deltanet_2_decoupling_erase_write_linear_attention | paper_analysis | memory_mechanism, llm, reasoning, long_term_memory, reinforce_learning | nvidia, deltanet | 100 | | | Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention | [object Object] | [object Object] | md | | |
| 20260528_CORE_contrastive_reflection_reasoning | paper_analysis | reasoning, memory_mechanism, self_evolving_agents, contrastive_reflection, reinforce_learning, cognitive_science | stanford_iris_lab, grpo, memgpt | 100 | | | CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning | [object Object] | [object Object] | md | | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key, MemLineage: Lineage-Guided Enforcement for LLM Agent Memory, AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment |
| 20260604_language_models_need_sleep_self_modify_consolidate_memories | paper_analysis | memory_mechanism, self_evolving_agents, continual_learning, reinforce_learning, reasoning, llm | google_brain, knowledge_seeding, sleep_paradigm | 100 | | | Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories | [object Object] | [object Object] | md | | memory_os_of_ai_agent, hipporag_neurobiologically_inspired_long_term_memory, muse_autoskill_self_evolving_skill_memory |
| 20260604_pretraining_recurrent_networks_without_recurrence | paper_analysis | memory_mechanism, recurrent_neural_networks, llm, reinforce_learning, neuro_science | mit, supervised_memory_training, backpropagation_through_time | 100 | | | Pretraining Recurrent Networks without Recurrence | [object Object] | [object Object] | md | | |
| 20260609_socratic_swe_self_evolving_coding_agents_trace_derived_skills | paper_analysis | self_evolving_agents, memory_mechanism, skill_curation, reinforce_learning, reasoning | sjtu, swebench_verified, terminalbench_2, grpo, skillos | 100 | | | Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills | [object Object] | [object Object] | md | | skillos_learning_skill_curation_self_evolving_agents, skillopt_executive_strategy_self_evolving_agent_skills, aflow_automating_agentic_workflow_generation, scaling_harness_agentic_ai, meme_multi_entity_evolving_memory_evaluation |
| 20260611_trace_unified_rollout_budget_allocation_agentic_rl | paper_analysis | reinforce_learning, reasoning, agent_architecture, test_time_scaling, llm | tsinghua_university, hotpotqa, bfcl_v3, monte_carlo_tree_search | 100 | | | TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning | [object Object] | [object Object] | md | | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_and_reasoning_frontier_survey_2026 |
| 20260624_world_models_in_pieces_structural_certification | paper_analysis | reinforce_learning, reasoning, agent_architecture, evaluation, foundation_agents | cuhk_shenzhen | 100 | | | World Models in Pieces: Structural Certification for General Agents | [object Object] | [object Object] | md | | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_characterization_system_implications |
| 20260626_progress_advantage_llm_agents | paper_analysis | reinforce_learning, reward_modeling, test_time_scaling, agent_architecture, llm, reasoning | university_of_wisconsin_madison, argonne_national_laboratory | 100 | | | Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents | [object Object] | [object Object] | md | | rl_long_horizon_reasoning_llm_expressiveness, vector_policy_optimization_vpo, trace_unified_rollout_budget_allocation_agentic_rl |
| 20260627_joint_learning_experiential_rules_policies_llm_agents | paper_analysis | memory_mechanism, reinforce_learning, self_evolving_agents, reasoning, agent_architecture | alfworld, grpo, sun_yat_sen_university | 100 | | | Joint Learning of Experiential Rules and Policies for Large Language Model Agents | [object Object] | [object Object] | md | | unlocking_working_memory_latent_reasoning, memory_os_of_ai_agent, memskill_learning_evolving_memory_skills |
| 20260629_memory_r1_enhancing_llm_agents_manage_utilize_memories_rl | paper_analysis | agent_architecture, memory_mechanism, reinforce_learning, llm, long_term_memory | lmu_munich, mem0, locomo_bench, grpo | 100 | | | Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning | [object Object] | [object Object] | md | | memgpt_towards_llms_as_operating_systems, muse_autoskill_self_evolving_skill_memory, gems_agent_native_multimodal_generation_memory_skills, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications |
| 20260701_severa_verified_synthesis_self_evolving_agents | paper_analysis | agent_architecture, self_evolving_agents, reinforce_learning, reasoning, agent_security | university_of_illinois_urbana_champaign, grpo, tau_squared_bench | 100 | | | SEVerA: Verified Synthesis of Self-Evolving Agents | [object Object] | [object Object] | md | | skillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, trace_unified_rollout_budget_agentic_rl, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications |
| 20260706_demopsd_disagreement_modulated_policy_self_distillation | paper_analysis | reinforce_learning, reasoning, llm, self_evolving_agents | grpo, sciknoweval, kl_distillation, gpqa | 100 | | | DemoPSD: Disagreement-Modulated Policy Self-Distillation | [object Object] | [object Object] | md | | rl_long_horizon_reasoning, vector_policy_optimization |
| 20260707_evopolicygym_evaluating_autonomous_policy_evolution | paper_analysis | embodied_ai, reinforce_learning, self_evolving_agents, evaluation, agent_architecture | cuhk_shenzhen, sjtu, tsinghua_university, ustc | 100 | | | EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments | [object Object] | [object Object] | md | | demopsd_disagreement_modulated_policy_self_distillation, llm_agents_social_structure_latent_objective_emergence, recontext_recursive_evidence_replay, evolvenav_proactive_preflection_self_evolving_memory |
| 20260708_adacurl_adaptive_curriculum_rl | paper_analysis | reinforce_learning, llm, reasoning, agent_architecture, evaluation | amap | 100 | | | AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting | [object Object] | [object Object] | md | | |
| 20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domains | paper_analysis | reinforce_learning, reward_modeling, llm, reasoning, evaluation | scale_ai | 100 | | | Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains | [object Object] | [object Object] | md | | |
| 20260723_copy_less_ground_more_evidence_aware_rl | paper_analysis | reasoning, reinforce_learning, reward_modeling, context_engineering, llm | pku, alibaba_group, gear, gspo, ruler_benchmark | 100 | | | Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning | [object Object] | [object Object] | md | | 20260703_self_evolving_world_models_llm_agent_planning, 20260719_searchos_v1_open_domain_information_seeking_agent_collaboration |
| 20260731_memrl_self_evolving_agents_runtime_rl_episodic_memory | paper_analysis | memory_mechanism, reinforce_learning, agent_architecture, llm, reasoning, long_term_memory | sjtu, xidian_university, nus, shanghai_innovation_institute, memtensor | 100 | | | MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory | [object Object] | [object Object] | md | | |