| 20260509_skillos_learning_skill_curation_self_evolving_agents | paper_analysis | agent_architecture, memory_mechanism, reinforce_learning, self_evolving_agents, skill_curation | google_cloud_ai_research, grpo, skillos, alfworld | 100 | | | SkillOS: Learning Skill Curation for Self-Evolving Agents | self_evolving_agents_survey_asi, memory_os_of_ai_agent, agentevolver_self_evolving_agent, reasoningbank_scaling_agent_self_evolving_reasoning_memory | [object Object] | [object Object] | md | |
| 20260520_amr_sd_token_level_credit_assignment | paper_analysis | reinforce_learning, reasoning, llm, reward_modeling, self_evolving_agents | grpo, dapo, meituan, sciknoweval | 100 | | | AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment | rl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents | [object Object] | [object Object] | md | |
| 20260524_vector_policy_optimization_vpo | paper_analysis | reinforce_learning, llm, test_time_scaling, reasoning, reward_modeling | mit, grpo | 100 | | | Vector Policy Optimization: Training for Diversity Improves Test-Time Search | | [object Object] | [object Object] | md | |
| 20260528_CORE_contrastive_reflection_reasoning | paper_analysis | reasoning, memory_mechanism, self_evolving_agents, contrastive_reflection, reinforce_learning, cognitive_science | stanford_iris_lab, grpo, memgpt | 100 | | | CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key, MemLineage: Lineage-Guided Enforcement for LLM Agent Memory, AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment | [object Object] | [object Object] | md | |
| 20260609_socratic_swe_self_evolving_coding_agents_trace_derived_skills | paper_analysis | self_evolving_agents, memory_mechanism, skill_curation, reinforce_learning, reasoning | sjtu, swebench_verified, terminalbench_2, grpo, skillos | 100 | | | Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills | skillos_learning_skill_curation_self_evolving_agents, skillopt_executive_strategy_self_evolving_agent_skills, aflow_automating_agentic_workflow_generation, scaling_harness_agentic_ai, meme_multi_entity_evolving_memory_evaluation | [object Object] | [object Object] | md | |
| 20260627_joint_learning_experiential_rules_policies_llm_agents | paper_analysis | memory_mechanism, reinforce_learning, self_evolving_agents, reasoning, agent_architecture | alfworld, grpo, sun_yat_sen_university | 100 | | | Joint Learning of Experiential Rules and Policies for Large Language Model Agents | unlocking_working_memory_latent_reasoning, memory_os_of_ai_agent, memskill_learning_evolving_memory_skills | [object Object] | [object Object] | md | |
| 20260629_memory_r1_enhancing_llm_agents_manage_utilize_memories_rl | paper_analysis | agent_architecture, memory_mechanism, reinforce_learning, llm, long_term_memory | lmu_munich, mem0, locomo_bench, grpo | 100 | | | Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning | memgpt_towards_llms_as_operating_systems, muse_autoskill_self_evolving_skill_memory, gems_agent_native_multimodal_generation_memory_skills, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications | [object Object] | [object Object] | md | |
| 20260701_severa_verified_synthesis_self_evolving_agents | paper_analysis | agent_architecture, self_evolving_agents, reinforce_learning, reasoning, agent_security | university_of_illinois_urbana_champaign, grpo, tau_squared_bench | 100 | | | SEVerA: Verified Synthesis of Self-Evolving Agents | skillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, trace_unified_rollout_budget_agentic_rl, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications | [object Object] | [object Object] | md | |
| 20260706_demopsd_disagreement_modulated_policy_self_distillation | paper_analysis | reinforce_learning, reasoning, llm, self_evolving_agents | grpo, sciknoweval, kl_distillation, gpqa | 100 | | | DemoPSD: Disagreement-Modulated Policy Self-Distillation | rl_long_horizon_reasoning, vector_policy_optimization | [object Object] | [object Object] | md | |