| 20260504_reasoningbank_scaling_agent_self_evolving_reasoning_memory | paper_analysis | self_evolving_agents, reasoning_memory, test_time_scaling, agent_architecture, memory_mechanism | google_cloud_ai_research, webarena, mind2web, swebench_verified | 100 | | | ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory | Self-Evolving Agents: A Survey | [object Object] | [object Object] | md | |
| 20260511_rl_long_horizon_reasoning_llm_expressiveness | paper_analysis | reinforce_learning, reasoning, llm, test_time_scaling, symbolic_reasoning | purdue_university, georgia_tech, scalelogic | 100 | | | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key | chain_of_thought_prompting_elicits_reasoning, reasoningbank_scaling_agent_self_evolving_reasoning_memory | [object Object] | [object Object] | md | |
| 20260524_vector_policy_optimization_vpo | paper_analysis | reinforce_learning, llm, test_time_scaling, reasoning, reward_modeling | mit, grpo | 100 | | | Vector Policy Optimization: Training for Diversity Improves Test-Time Search | | [object Object] | [object Object] | md | |
| 20260611_trace_unified_rollout_budget_allocation_agentic_rl | paper_analysis | reinforce_learning, reasoning, agent_architecture, test_time_scaling, llm | tsinghua_university, hotpotqa, bfcl_v3, monte_carlo_tree_search | 100 | | | TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_and_reasoning_frontier_survey_2026 | [object Object] | [object Object] | md | |
| 20260624_self_harness | paper_analysis | agent_architecture, self_evolving_agents, llm, reasoning, test_time_scaling | shanghai_ai_lab, terminalbench_2 | 100 | | | Self-Harness: Harnesses That Improve Themselves | | [object Object] | [object Object] | md | |
| 20260626_progress_advantage_llm_agents | paper_analysis | reinforce_learning, reward_modeling, test_time_scaling, agent_architecture, llm, reasoning | university_of_wisconsin_madison, argonne_national_laboratory | 100 | | | Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents | rl_long_horizon_reasoning_llm_expressiveness, vector_policy_optimization_vpo, trace_unified_rollout_budget_allocation_agentic_rl | [object Object] | [object Object] | md | |
| 20260708_solving_million_step_llm_zero_errors | paper_analysis | multi_agent_systems, reasoning, agent_architecture, llm, test_time_scaling | cognizant_ai_lab, ut_austin | 100 | | | Solving a Million-Step LLM Task with Zero Errors | | [object Object] | [object Object] | md | |