| 20260422_llm_semantic_reasoners_not_symbolic_reasoners | paper_analysis | llm, reasoning, cognitive_science, language_philosophy | pku, bigai, proofwriter | 100 | | | Large Language Models are In-Context Semantic Reasoners rather than Symbolic Reasoners | HippoRAG: Neurobiologically Inspired Long-Term Memory, M3-Agent: Multimodal Agent with Long-Term Memory, Memory OS of AI Agent | [object Object] | [object Object] | md | |
| 20260423_agentevolver_self_evolving_agent | paper_analysis | agent_architecture, reinforce_learning, llm, reasoning | tongyi_lab, appworld, bfcl_v3 | 100 | | | AgentEvolver: Towards Efficient Self-Evolving Agent System | | [object Object] | [object Object] | md | |
| 20260423_ai_reasoning_deep_learning_symbolic_neural | paper_analysis | reasoning, symbolic_reasoning, llm, knowledge_graph, neuro_science, cognitive_science | beihang_university | 100 | | | AI Reasoning in Deep Learning Era: From Symbolic AI to Neural Symbolic AI | attentive_reasoning_queries | [object Object] | [object Object] | md | |
| 20260423_attentive_reasoning_queries | paper_analysis | llm, agent_architecture, reasoning, cognitive_science, context_engineering | emcie_co | 100 | | | Attentive Reasoning Queries: Optimizing Instruction-Following in LLMs | ai_reasoning_deep_learning_symbolic_neural | [object Object] | [object Object] | md | |
| 20260423_meta_harness_end_to_end_optimization | paper_analysis | context_engineering, agent_architecture, llm, reasoning, rag | stanford_iris_lab, terminalbench_2, krafton | 100 | | | Meta-Harness: End-to-End Optimization of Model Harnesses | | [object Object] | [object Object] | md | |
| 20260424_context_engineering_2_overview | paper_analysis | context_engineering, llm, agent_architecture, reasoning | sjtu_gair_lab, tongyi_lab | 100 | | | context_engineering_2_overview | | [object Object] | [object Object] | md | |
| 20260424_llm_knowledge_graphs_opportunities_challenges | paper_analysis | knowledge_graph, llm, rag, reasoning | | 100 | | | llm_knowledge_graphs_opportunities_challenges | | [object Object] | [object Object] | md | |
| 20260425_chain_of_thought_prompting_elicits_reasoning | paper_analysis | llm, reasoning, cognitive_science, context_engineering, multi_hop_reasoning | google_brain, chain_of_thought, gsm8k | 100 | | | Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | | [object Object] | [object Object] | md | |
| 20260427_agentsquare_automatic_llm_agent_search | paper_analysis | agent_architecture, llm, reasoning, memory_mechanism, reinforce_learning | tsinghua_fib_lab, neural_architecture_search | 100 | | | AgentSquare: Automatic LLM Agent Search in Modular Design Space | | [object Object] | [object Object] | md | |
| 20260427_detecting_hallucinations_semantic_entropy | paper_analysis | llm, reasoning, agent_architecture, memory_mechanism, reinforce_learning | oatml_oxford, neural_architecture_search | 100 | | | Detecting hallucinations in large language models using semantic entropy | | [object Object] | [object Object] | md | |
| 20260428_rubrics_as_rewards | paper_analysis | reinforce_learning, llm, reasoning, reward_modeling | scale_ai | 95 | | | Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains | | [object Object] | [object Object] | md | |
| 20260429_disentangling_memory_reasoning_llm | paper_analysis | memory_mechanism, reasoning, llm, multi_hop_reasoning, chain_of_thought, cognitive_science | rutgers_university, ohio_state_university, ucsb, truthfulqa, strategyqa, commonsenseqa | 100 | | | Disentangling Memory and Reasoning Ability in Large Language Models | agent_workflow_memory | [object Object] | [object Object] | md | |
| 20260430_catastrophic_forgetting_implicit_inference | paper_analysis | catastrophic_forgetting, llm, memory_mechanism, reasoning, context_engineering | cmu, conjugate_prompting | 100 | | | Understanding Catastrophic Forgetting in Language Models via Implicit Inference | | [object Object] | [object Object] | md | |
| 20260503_aflow_automating_agentic_workflow_generation | paper_analysis | agent_architecture, llm, reasoning, workflow_optimization | deepwisdom, monte_carlo_tree_search, claw | 100 | | | AFLOW: Automating Agentic Workflow Generation | | [object Object] | [object Object] | md | |
| 20260506_openclaw_rl_train_any_agent_simply_by_talking | paper_analysis | reinforce_learning, agent_architecture, llm, self_evolving_agents, reasoning | gen_verse | 100 | | | OpenClaw-RL: Train Any Agent Simply by Talking | | [object Object] | [object Object] | md | |
| 20260511_rl_long_horizon_reasoning_llm_expressiveness | paper_analysis | reinforce_learning, reasoning, llm, test_time_scaling, symbolic_reasoning | purdue_university, georgia_tech, scalelogic | 100 | | | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key | chain_of_thought_prompting_elicits_reasoning, reasoningbank_scaling_agent_self_evolving_reasoning_memory | [object Object] | [object Object] | md | |
| 20260512_meme_multi_entity_evolving_memory_evaluation | paper_analysis | memory_mechanism, long_term_memory, agent_architecture, llm, reasoning, evaluation | kaist_ai, tuebingen_ai_center, naver_ai_lab | 100 | | | MEME: Multi-entity & Evolving Memory Evaluation | memory_os_of_ai_agent, memgpt_towards_llms_as_operating_systems, disentangling_memory_reasoning_llm | [object Object] | [object Object] | md | |
| 20260515_apwa_parallelizable_agentic_workflows | paper_analysis | multi_agent_systems, agent_architecture, distributed_systems, workflow_optimization, reasoning, llm | northeastern_university, apwa, ray, pii_300k, schemabench, summarybench | 95 | | | APWA: A Distributed Architecture for Parallelizable Agentic Workflows | scaling_large_language_model_multi_agent_collaboration, gasim_graph_accelerated_hybrid_social_simulation | [object Object] | [object Object] | md | |
| 20260515_harnessing_agentic_evolution | paper_analysis | agent_architecture, self_evolving_agents, reasoning, memory_mechanism, reinforce_learning, workflow_optimization | deepwisdom, hkust_gz, sjtu, tsinghua_university, nanyang_technological_university | 100 | | | harnessing_agentic_evolution | self_evolving_agents_survey_asi, skillos_learning_skill_curation_self_evolving_agents, aflow_automating_agentic_workflow_generation | [object Object] | [object Object] | md | |
| 20260517_heterogeneous_temporal_memory_governance_llm_persona | paper_analysis | memory_mechanism, llm, long_term_memory, rag, reasoning, persona_consistency | uestc, arpm | 100 | | | A Heterogeneous Temporal Memory Governance Framework for Long-Term LLM Persona Consistency | memgpt_towards_llms_as_operating_systems, meme_multi_entity_evolving_memory_evaluation | [object Object] | [object Object] | md | |
| 20260520_amr_sd_token_level_credit_assignment | paper_analysis | reinforce_learning, reasoning, llm, reward_modeling, self_evolving_agents | grpo, dapo, meituan, sciknoweval | 100 | | | AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment | rl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents | [object Object] | [object Object] | md | |
| 20260520_skillgenbench_benchmarking_skill_generation_llm_agents | paper_analysis | agent_architecture, evaluation, skill_curation, llm, reasoning | quanta_alpha, sjtu, pku, nus, xjtu, tsinghua_university, sufe, ntu, ucas | 100 | | | SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents | ontology_skill_analysis | [object Object] | [object Object] | md | |
| 20260521_methodology_selecting_composing_runtime_architecture_patterns_production_llm_agents | paper_analysis | agent_architecture, distributed_systems, multi_agent_systems, llm, reasoning | stanford_iris_lab, claw | 100 | | | A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents | agent_memory_and_reasoning_frontier_survey_2026, memgpt_towards_llms_as_operating_systems, memlineage_lineage_guided_llm_agent_memory | [object Object] | [object Object] | md | |
| 20260524_vector_policy_optimization_vpo | paper_analysis | reinforce_learning, llm, test_time_scaling, reasoning, reward_modeling | mit, grpo | 100 | | | Vector Policy Optimization: Training for Diversity Improves Test-Time Search | | [object Object] | [object Object] | md | |
| 20260525_gated_deltanet_2_decoupling_erase_write_linear_attention | paper_analysis | memory_mechanism, llm, reasoning, long_term_memory, reinforce_learning | nvidia, deltanet | 100 | | | Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention | | [object Object] | [object Object] | md | |
| 20260526_skillopt_executive_strategy_self_evolving_agent_skills | paper_analysis | self_evolving_agents, skill_curation, agent_architecture, reasoning, workflow_optimization | microsoft_research_asia, sjtu | 100 | | | SkillOpt: Executive Strategy for Self-Evolving Agent Skills | skillos, moss, harnessing_agentic_evolution, self_evolving_agents_survey | [object Object] | [object Object] | md | |
| 20260527_scaling_harness_agentic_ai | paper_analysis | agent_architecture, memory_mechanism, reasoning, self_evolving_agents, multi_agent_systems | uc_berkeley, cheetahclaws | 100 | | | From Model Scaling to System Scaling: Scaling the Harness in Agentic AI | MemLineage, MemGPT, SkillOS, CalMem | [object Object] | [object Object] | md | |
| 20260528_CORE_contrastive_reflection_reasoning | paper_analysis | reasoning, memory_mechanism, self_evolving_agents, contrastive_reflection, reinforce_learning, cognitive_science | stanford_iris_lab, grpo, memgpt | 100 | | | CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning | Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key, MemLineage: Lineage-Guided Enforcement for LLM Agent Memory, AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment | [object Object] | [object Object] | md | |
| 20260528_unlocking_working_memory_latent_reasoning | paper_analysis | memory_mechanism, reasoning, llm, reasoning_memory, context_engineering | lit_ai_lab, nxai_gmbh, gsm8k | 100 | | | Unlocking the Working Memory of Large Language Models for Latent Reasoning | chain_of_thought_prompting_elicits_reasoning, memgpt_towards_llms_as_operating_systems, agent_memory_and_reasoning_frontier_survey_2026 | [object Object] | [object Object] | md | |
| 20260529_stop_comparing_llm_agents_without_disclosing_harness | paper_analysis | agent_architecture, evaluation, llm, multi_agent_systems, reasoning | tulane_university, rutgers_university, virginia_tech | 100 | | | Stop Comparing LLM Agents Without Disclosing the Harness | Stop Comparing LLM Agents Without Disclosing the Harness | [object Object] | [object Object] | md | |
| 20260601_locally_coherent_globally_incoherent_multi_component_llm_agents | paper_analysis | multi_agent_systems, reasoning, agent_architecture, evaluation | princeton_university, paleka | 100 | | | Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents | scaling_large_language_model_multi_agent_collaboration, gasim_graph_accelerated_hybrid_social_simulation, harnessing_agentic_evolution, scaling_harness_agentic_ai | [object Object] | [object Object] | md | |
| 20260602_lintree_explicitly_structured_search_histories | paper_analysis|paper_analysis | reasoning, llm, agent_architecture, memory_mechanism | nus, oatml_oxford | 100 | | | LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories | chain_of_thought_prompting_elicits_reasoning | [object Object] | [object Object] | md | |
| 20260604_language_models_need_sleep_self_modify_consolidate_memories | paper_analysis | memory_mechanism, self_evolving_agents, continual_learning, reinforce_learning, reasoning, llm | google_brain, knowledge_seeding, sleep_paradigm | 100 | | | Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories | memory_os_of_ai_agent, hipporag_neurobiologically_inspired_long_term_memory, muse_autoskill_self_evolving_skill_memory | [object Object] | [object Object] | md | |
| 20260605_aip_graph_representation_for_learning_and_governing_agent_skills | paper_analysis | agent_architecture, skill_curation, self_evolving_agents, reasoning, workflow_optimization | skillsbench, agent_instruction_protocol | 95 | | | AIP: A Graph Representation for Learning and Governing Agent Skills | skillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, aflow_automating_agentic_workflow_generation, apwa_parallelizable_agentic_workflows | [object Object] | [object Object] | md | |
| 20260609_socratic_swe_self_evolving_coding_agents_trace_derived_skills | paper_analysis | self_evolving_agents, memory_mechanism, skill_curation, reinforce_learning, reasoning | sjtu, swebench_verified, terminalbench_2, grpo, skillos | 100 | | | Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills | skillos_learning_skill_curation_self_evolving_agents, skillopt_executive_strategy_self_evolving_agent_skills, aflow_automating_agentic_workflow_generation, scaling_harness_agentic_ai, meme_multi_entity_evolving_memory_evaluation | [object Object] | [object Object] | md | |
| 20260610_searchswarm_delegation_intelligence_agentic_llm_long_horizon_research | paper_analysis | agent_architecture, multi_agent_systems, reasoning, workflow_optimization, reasoning_memory | tsinghua_university, pku, ant_group, browsecomp, searchswarm | 100 | | | SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research | lintree_explicitly_structured_search_histories, aip_graph_representation_for_learning_and_governing_agent_skills, aflow_automating_agentic_workflow_generation, apwa_parallelizable_agentic_workflows, scaling_large_language_model_multi_agent_collaboration | [object Object] | [object Object] | md | |
| 20260611_trace_unified_rollout_budget_allocation_agentic_rl | paper_analysis | reinforce_learning, reasoning, agent_architecture, test_time_scaling, llm | tsinghua_university, hotpotqa, bfcl_v3, monte_carlo_tree_search | 100 | | | TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_and_reasoning_frontier_survey_2026 | [object Object] | [object Object] | md | |
| 20260618_fixed_point_reasoners_stable_adaptive_deep_looped_transformers | paper_analysis | reasoning, llm, recurrent_neural_networks, memory_mechanism, agent_architecture | tuebingen_ai_center, eth_zurich, liquidai, max_planck_institute | 100 | | | Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers | | [object Object] | [object Object] | md | |
| 20260622_how_transparent_is_diffusiongemma | paper_analysis | llm, reasoning, agent_security, evaluation, interpretability | google_brain | 100 | | | How Transparent is DiffusionGemma? | detecting_hallucinations_semantic_entropy | [object Object] | [object Object] | md | |
| 20260624_self_harness | paper_analysis | agent_architecture, self_evolving_agents, llm, reasoning, test_time_scaling | shanghai_ai_lab, terminalbench_2 | 100 | | | Self-Harness: Harnesses That Improve Themselves | | [object Object] | [object Object] | md | |
| 20260624_world_models_in_pieces_structural_certification | paper_analysis | reinforce_learning, reasoning, agent_architecture, evaluation, foundation_agents | cuhk_shenzhen | 100 | | | World Models in Pieces: Structural Certification for General Agents | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_characterization_system_implications | [object Object] | [object Object] | md | |
| 20260626_progress_advantage_llm_agents | paper_analysis | reinforce_learning, reward_modeling, test_time_scaling, agent_architecture, llm, reasoning | university_of_wisconsin_madison, argonne_national_laboratory | 100 | | | Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents | rl_long_horizon_reasoning_llm_expressiveness, vector_policy_optimization_vpo, trace_unified_rollout_budget_allocation_agentic_rl | [object Object] | [object Object] | md | |
| 20260627_joint_learning_experiential_rules_policies_llm_agents | paper_analysis | memory_mechanism, reinforce_learning, self_evolving_agents, reasoning, agent_architecture | alfworld, grpo, sun_yat_sen_university | 100 | | | Joint Learning of Experiential Rules and Policies for Large Language Model Agents | unlocking_working_memory_latent_reasoning, memory_os_of_ai_agent, memskill_learning_evolving_memory_skills | [object Object] | [object Object] | md | |
| 20260628_omniact_omnimodal_embodied_agents | paper_analysis | embodied_ai, agent_architecture, long_term_memory, multimodal, robotics, reasoning | fudan_university, memgpt | 100 | | | Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy | memgpt_towards_llms_as_operating_systems, memory_os_of_ai_agent, reca_integrated_acceleration_cooperative_embodied_agents | [object Object] | [object Object] | md | |
| 20260629_carve_content_aware_recurrent_value_efficiency | paper_analysis | memory_mechanism, recurrent_neural_networks, llm, reasoning, linear_attention | carve, deltanet | 100 | | | CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention | Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention | [object Object] | [object Object] | md | |
| 20260630_from_tokens_to_states_llms_world_models | paper_analysis | llm, reasoning, world_model, memory_mechanism, agent_architecture | jepa, othello_gpt | 100 | | | From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond | memory_os_of_ai_agent, rl_long_horizon_reasoning_llm_expressiveness | [object Object] | [object Object] | md | |
| 20260701_severa_verified_synthesis_self_evolving_agents | paper_analysis | agent_architecture, self_evolving_agents, reinforce_learning, reasoning, agent_security | university_of_illinois_urbana_champaign, grpo, tau_squared_bench | 100 | | | SEVerA: Verified Synthesis of Self-Evolving Agents | skillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, trace_unified_rollout_budget_agentic_rl, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications | [object Object] | [object Object] | md | |
| 20260703_self_evolving_world_models_llm_agent_planning | paper_analysis | world_model, self_evolving_agents, memory_mechanism, embodied_ai, reasoning | nus, alfworld, deepmind | 100 | | | Self-Evolving World Models for LLM Agent Planning | from_tokens_to_states_llms_world_models, world_models_in_pieces_structural_certification, memory_r1_enhancing_llm_agents_manage_utilize_memories_rl, stop_comparing_llm_agents_without_disclosing_harness | [object Object] | [object Object] | md | |
| 20260705_llm_agents_social_structure_latent_objective_emergence | paper_analysis | multi_agent_systems, social_simulation, agent_architecture, reasoning, llm, evaluation | cmu, dual_channel_debate, llm_agora | 100 | | | What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates | self_harness, contagion_networks_evaluator_bias_multi_agent_llm | [object Object] | [object Object] | md | |
| 20260706_demopsd_disagreement_modulated_policy_self_distillation | paper_analysis | reinforce_learning, reasoning, llm, self_evolving_agents | grpo, sciknoweval, kl_distillation, gpqa | 100 | | | DemoPSD: Disagreement-Modulated Policy Self-Distillation | rl_long_horizon_reasoning, vector_policy_optimization | [object Object] | [object Object] | md | |
| 20260708_adacurl_adaptive_curriculum_rl | paper_analysis | reinforce_learning, llm, reasoning, agent_architecture, evaluation | amap | 100 | | | AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting | | [object Object] | [object Object] | md | |
| 20260708_solving_million_step_llm_zero_errors | paper_analysis | multi_agent_systems, reasoning, agent_architecture, llm, test_time_scaling | cognizant_ai_lab, ut_austin | 100 | | | Solving a Million-Step LLM Task with Zero Errors | | [object Object] | [object Object] | md | |
| 20260710_let_the_agent_search_tkgqa | paper_analysis | knowledge_graph, multi_hop_reasoning, llm, agent_architecture, reasoning | at2qa, multi_tq, timeline_cronquestion, timeline_icews_actor | 100 | | | Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering | | [object Object] | [object Object] | md | |
| 20260713_metacognition_in_llms_foundations_progress_opportunities | paper_analysis, domain_survey | llm, reasoning, cognitive_science, memory_mechanism, agent_architecture | yale_university, uc_irvine | 100 | | | Metacognition in LLMs: Foundations, Progress, and Opportunities | chain_of_thought_prompting_elicits_reasoning, ai_reasoning_deep_learning_symbolic_neural, memory_os_of_ai_agent | [object Object] | [object Object] | md | |
| 20260714_comprehensive_survey_knowledge_graph_reasoning_approaches_applications | paper_analysis | knowledge_graph, multi_hop_reasoning, reasoning, llm, survey | | 100 | | | A Comprehensive Survey of Knowledge Graph Reasoning: Approaches and Applications | | [object Object] | [object Object] | md | |
| 20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domains | paper_analysis | reinforce_learning, reward_modeling, llm, reasoning, evaluation | scale_ai | 100 | | | Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains | | [object Object] | [object Object] | md | |
| 20260716_prefill_as_service_kv_cache_cross_datacenter | paper_analysis | llm_serving, kv_cache, distributed_systems, multi_agent_systems, reasoning | moonshot_ai, tsinghua_university | 100 | | | Prefill-as-a-Service: KV Cache of Next-Generation Models Could Go Cross-Datacenter | 20260714_comprehensive_survey_knowledge_graph_reasoning_approaches_applications, 20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domains | [object Object] | [object Object] | md | |
| 20260717_experience_memory_graph_one_shot_error_correction_for_agents | paper_analysis | memory_mechanism, agent_architecture, reasoning, long_term_memory, knowledge_graph | alfworld, scienceworld, uestc | 100 | | | Experience Memory Graph: One-Shot Error Correction for Agents | memlineage_lineage_guided_llm_agent_memory, shared_selective_persistent_memory_agentic_llm, agent_memory_and_reasoning_frontier_survey_2026 | [object Object] | [object Object] | md | |
| 20260719_t2mlr_transformer_temporal_middle_layer_recurrence | paper_analysis | reasoning, llm, memory_mechanism, recurrent_neural_networks, reasoning_memory | princeton_university, chain_of_thought, backpropagation_through_time, gsm8k, hotpotqa | 100 | | | T²MLR: Transformer with Temporal Middle-Layer Recurrence | searchos_v1_open_domain_information_seeking_agent_collaboration, chain_of_thought_prompting_elicits_reasoning, memory_r1_enhancing_llm_agents_manage_utilize_memories_rl | [object Object] | [object Object] | md | |
| 20260722_supra_cognitive_modes_routed_agent_memory | paper_analysis | agent_architecture, memory_mechanism, long_term_memory, reasoning, multi_hop_reasoning | supra_research, locomo_bench, memory_agent_bench, long_mem_eval | 100 | | | Supra Cognitive Modes: A Routed Architecture for Agent Memory | memory_os_of_ai_agent, memgpt_towards_llms_as_operating_systems, shared_selective_persistent_memory_agentic_llm | [object Object] | [object Object] | md | |
| 20260723_copy_less_ground_more_evidence_aware_rl | paper_analysis | reasoning, reinforce_learning, reward_modeling, context_engineering, llm | pku, alibaba_group, gear, gspo, ruler_benchmark | 100 | | | Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning | 20260703_self_evolving_world_models_llm_agent_planning, 20260719_searchos_v1_open_domain_information_seeking_agent_collaboration | [object Object] | [object Object] | md | |
| 20260724_pro_long_programmatic_memory_long_horizon_reasoning | paper_analysis | memory_mechanism, agent_architecture, reasoning, long_term_memory, coding_agents, continual_learning | duke_university, arc_agi, chain_of_thought | 100 | | | PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning | | [object Object] | [object Object] | md | |
| 20260725_evolving_commitments_self_adaptive_sts | paper_analysis | multi_agent_systems, agent_architecture, software_engineering, self_evolving_agents, reasoning | fudan_university, the_open_university, university_of_trento | 100 | | | Evolving Commitments for Self-Adaptive Socio-Technical Systems | | [object Object] | [object Object] | md | |
| 20260725_monitoring_diagnosing_software_requirements | paper_analysis | software_engineering, reasoning, symbolic_reasoning, knowledge_graph, agent_architecture | university_of_toronto, the_open_university | 100 | | | Monitoring and Diagnosing Software Requirements | | [object Object] | [object Object] | md | |
| 20260726_agentic_context_management_memory_cost_lifecycle_architecture | paper_analysis | agent_architecture, memory_mechanism, context_engineering, reasoning, long_term_memory | maximem, long_mem_eval, locomo_bench | 100 | | | Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems | memory_os_of_ai_agent, agent_memory_and_reasoning_frontier_survey_2026, memory_in_the_age_of_ai_agents_a_survey, memgpt_towards_llms_as_operating_systems, context_engineering_2_overview | [object Object] | [object Object] | md | |
| 20260731_memrl_self_evolving_agents_runtime_rl_episodic_memory | paper_analysis | memory_mechanism, reinforce_learning, agent_architecture, llm, reasoning, long_term_memory | sjtu, xidian_university, nus, shanghai_innovation_institute, memtensor | 100 | | | MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory | | [object Object] | [object Object] | md | |