Reinforcement learning theory and applications

表格 30 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleNavigationOrderEleventyTemplateEngineOverrideCreatedRelatedNotes
20260423_agentevolver_self_evolving_agentpaper_analysisagent_architecture, reinforce_learning, llm, reasoningtongyi_lab, appworld, bfcl_v3100AgentEvolver: Towards Efficient Self-Evolving Agent System[object Object][object Object]md
20260427_agentsquare_automatic_llm_agent_searchpaper_analysisagent_architecture, llm, reasoning, memory_mechanism, reinforce_learningtsinghua_fib_lab, neural_architecture_search100AgentSquare: Automatic LLM Agent Search in Modular Design Space[object Object][object Object]md
20260427_detecting_hallucinations_semantic_entropypaper_analysisllm, reasoning, agent_architecture, memory_mechanism, reinforce_learningoatml_oxford, neural_architecture_search100Detecting hallucinations in large language models using semantic entropy[object Object][object Object]md
20260428_navigating_to_objects_in_the_real_worldpaper_analysisembodied_ai, reinforce_learning, sim_to_real, semantic_navigationmeta_ai, habitat_simulator100Navigating to objects in the real world[object Object][object Object]md
20260428_rubrics_as_rewardspaper_analysisreinforce_learning, llm, reasoning, reward_modelingscale_ai95Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains[object Object][object Object]md
20260430_hallucination_survey_llmpaper_analysishallucination, llm, survey, rag, reinforce_learningharbin_institute_of_technology100A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions[object Object][object Object]mdcatastrophic_forgetting_implicit_inference
20260506_openclaw_rl_train_any_agent_simply_by_talkingpaper_analysisreinforce_learning, agent_architecture, llm, self_evolving_agents, reasoninggen_verse100OpenClaw-RL: Train Any Agent Simply by Talking[object Object][object Object]md
20260509_skillos_learning_skill_curation_self_evolving_agentspaper_analysisagent_architecture, memory_mechanism, reinforce_learning, self_evolving_agents, skill_curationgoogle_cloud_ai_research, grpo, skillos, alfworld100SkillOS: Learning Skill Curation for Self-Evolving Agents[object Object][object Object]mdself_evolving_agents_survey_asi, memory_os_of_ai_agent, agentevolver_self_evolving_agent, reasoningbank_scaling_agent_self_evolving_reasoning_memory
20260511_rl_long_horizon_reasoning_llm_expressivenesspaper_analysisreinforce_learning, reasoning, llm, test_time_scaling, symbolic_reasoningpurdue_university, georgia_tech, scalelogic100Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key[object Object][object Object]mdchain_of_thought_prompting_elicits_reasoning, reasoningbank_scaling_agent_self_evolving_reasoning_memory
20260512_learning_fast_slow_llms_adapt_continuallypaper_analysiscatastrophic_forgetting, continual_learning, memory_mechanism, reinforce_learning, llmuc_berkeley100Learning, Fast and Slow: Towards LLMs That Adapt Continually[object Object][object Object]md20260504_reasoningbank_scaling_agent_self_evolving_reasoning_memory, 20260511_rl_long_horizon_reasoning_llm_expressiveness
20260515_harnessing_agentic_evolutionpaper_analysisagent_architecture, self_evolving_agents, reasoning, memory_mechanism, reinforce_learning, workflow_optimizationdeepwisdom, hkust_gz, sjtu, tsinghua_university, nanyang_technological_university100harnessing_agentic_evolution[object Object][object Object]mdself_evolving_agents_survey_asi, skillos_learning_skill_curation_self_evolving_agents, aflow_automating_agentic_workflow_generation
20260520_amr_sd_token_level_credit_assignmentpaper_analysisreinforce_learning, reasoning, llm, reward_modeling, self_evolving_agentsgrpo, dapo, meituan, sciknoweval100AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment[object Object][object Object]mdrl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents
20260524_vector_policy_optimization_vpopaper_analysisreinforce_learning, llm, test_time_scaling, reasoning, reward_modelingmit, grpo100Vector Policy Optimization: Training for Diversity Improves Test-Time Search[object Object][object Object]md
20260525_gated_deltanet_2_decoupling_erase_write_linear_attentionpaper_analysismemory_mechanism, llm, reasoning, long_term_memory, reinforce_learningnvidia, deltanet100Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention[object Object][object Object]md
20260528_CORE_contrastive_reflection_reasoningpaper_analysisreasoning, memory_mechanism, self_evolving_agents, contrastive_reflection, reinforce_learning, cognitive_sciencestanford_iris_lab, grpo, memgpt100CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning[object Object][object Object]mdCan RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key, MemLineage: Lineage-Guided Enforcement for LLM Agent Memory, AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
20260604_language_models_need_sleep_self_modify_consolidate_memoriespaper_analysismemory_mechanism, self_evolving_agents, continual_learning, reinforce_learning, reasoning, llmgoogle_brain, knowledge_seeding, sleep_paradigm100Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories[object Object][object Object]mdmemory_os_of_ai_agent, hipporag_neurobiologically_inspired_long_term_memory, muse_autoskill_self_evolving_skill_memory
20260604_pretraining_recurrent_networks_without_recurrencepaper_analysismemory_mechanism, recurrent_neural_networks, llm, reinforce_learning, neuro_sciencemit, supervised_memory_training, backpropagation_through_time100Pretraining Recurrent Networks without Recurrence[object Object][object Object]md
20260609_socratic_swe_self_evolving_coding_agents_trace_derived_skillspaper_analysisself_evolving_agents, memory_mechanism, skill_curation, reinforce_learning, reasoningsjtu, swebench_verified, terminalbench_2, grpo, skillos100Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills[object Object][object Object]mdskillos_learning_skill_curation_self_evolving_agents, skillopt_executive_strategy_self_evolving_agent_skills, aflow_automating_agentic_workflow_generation, scaling_harness_agentic_ai, meme_multi_entity_evolving_memory_evaluation
20260611_trace_unified_rollout_budget_allocation_agentic_rlpaper_analysisreinforce_learning, reasoning, agent_architecture, test_time_scaling, llmtsinghua_university, hotpotqa, bfcl_v3, monte_carlo_tree_search100TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning[object Object][object Object]mdrl_long_horizon_reasoning_llm_expressiveness, agent_memory_and_reasoning_frontier_survey_2026
20260624_world_models_in_pieces_structural_certificationpaper_analysisreinforce_learning, reasoning, agent_architecture, evaluation, foundation_agentscuhk_shenzhen100World Models in Pieces: Structural Certification for General Agents[object Object][object Object]mdrl_long_horizon_reasoning_llm_expressiveness, agent_memory_characterization_system_implications
20260626_progress_advantage_llm_agentspaper_analysisreinforce_learning, reward_modeling, test_time_scaling, agent_architecture, llm, reasoninguniversity_of_wisconsin_madison, argonne_national_laboratory100Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents[object Object][object Object]mdrl_long_horizon_reasoning_llm_expressiveness, vector_policy_optimization_vpo, trace_unified_rollout_budget_allocation_agentic_rl
20260627_joint_learning_experiential_rules_policies_llm_agentspaper_analysismemory_mechanism, reinforce_learning, self_evolving_agents, reasoning, agent_architecturealfworld, grpo, sun_yat_sen_university100Joint Learning of Experiential Rules and Policies for Large Language Model Agents[object Object][object Object]mdunlocking_working_memory_latent_reasoning, memory_os_of_ai_agent, memskill_learning_evolving_memory_skills
20260629_memory_r1_enhancing_llm_agents_manage_utilize_memories_rlpaper_analysisagent_architecture, memory_mechanism, reinforce_learning, llm, long_term_memorylmu_munich, mem0, locomo_bench, grpo100Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning[object Object][object Object]mdmemgpt_towards_llms_as_operating_systems, muse_autoskill_self_evolving_skill_memory, gems_agent_native_multimodal_generation_memory_skills, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications
20260701_severa_verified_synthesis_self_evolving_agentspaper_analysisagent_architecture, self_evolving_agents, reinforce_learning, reasoning, agent_securityuniversity_of_illinois_urbana_champaign, grpo, tau_squared_bench100SEVerA: Verified Synthesis of Self-Evolving Agents[object Object][object Object]mdskillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, trace_unified_rollout_budget_agentic_rl, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications
20260706_demopsd_disagreement_modulated_policy_self_distillationpaper_analysisreinforce_learning, reasoning, llm, self_evolving_agentsgrpo, sciknoweval, kl_distillation, gpqa100DemoPSD: Disagreement-Modulated Policy Self-Distillation[object Object][object Object]mdrl_long_horizon_reasoning, vector_policy_optimization
20260707_evopolicygym_evaluating_autonomous_policy_evolutionpaper_analysisembodied_ai, reinforce_learning, self_evolving_agents, evaluation, agent_architecturecuhk_shenzhen, sjtu, tsinghua_university, ustc100EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments[object Object][object Object]mddemopsd_disagreement_modulated_policy_self_distillation, llm_agents_social_structure_latent_objective_emergence, recontext_recursive_evidence_replay, evolvenav_proactive_preflection_self_evolving_memory
20260708_adacurl_adaptive_curriculum_rlpaper_analysisreinforce_learning, llm, reasoning, agent_architecture, evaluationamap100AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting[object Object][object Object]md
20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domainspaper_analysisreinforce_learning, reward_modeling, llm, reasoning, evaluationscale_ai100Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains[object Object][object Object]md
20260723_copy_less_ground_more_evidence_aware_rlpaper_analysisreasoning, reinforce_learning, reward_modeling, context_engineering, llmpku, alibaba_group, gear, gspo, ruler_benchmark100Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning[object Object][object Object]md20260703_self_evolving_world_models_llm_agent_planning, 20260719_searchos_v1_open_domain_information_seeking_agent_collaboration
20260731_memrl_self_evolving_agents_runtime_rl_episodic_memorypaper_analysismemory_mechanism, reinforce_learning, agent_architecture, llm, reasoning, long_term_memorysjtu, xidian_university, nus, shanghai_innovation_institute, memtensor100MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory[object Object][object Object]md
Powered by Forestry.md