Grouped Reward Policy Optimization, a reinforcement learning algorithm for training language models with group-based reward estimation.

表格 9 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleRelatedNotesNavigationOrderEleventyTemplateEngineOverrideCreated
20260509_skillos_learning_skill_curation_self_evolving_agentspaper_analysisagent_architecture, memory_mechanism, reinforce_learning, self_evolving_agents, skill_curationgoogle_cloud_ai_research, grpo, skillos, alfworld100SkillOS: Learning Skill Curation for Self-Evolving Agentsself_evolving_agents_survey_asi, memory_os_of_ai_agent, agentevolver_self_evolving_agent, reasoningbank_scaling_agent_self_evolving_reasoning_memory[object Object][object Object]md
20260520_amr_sd_token_level_credit_assignmentpaper_analysisreinforce_learning, reasoning, llm, reward_modeling, self_evolving_agentsgrpo, dapo, meituan, sciknoweval100AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignmentrl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents[object Object][object Object]md
20260524_vector_policy_optimization_vpopaper_analysisreinforce_learning, llm, test_time_scaling, reasoning, reward_modelingmit, grpo100Vector Policy Optimization: Training for Diversity Improves Test-Time Search[object Object][object Object]md
20260528_CORE_contrastive_reflection_reasoningpaper_analysisreasoning, memory_mechanism, self_evolving_agents, contrastive_reflection, reinforce_learning, cognitive_sciencestanford_iris_lab, grpo, memgpt100CORE: Contrastive Reflection Enables Rapid Improvements in ReasoningCan RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key, MemLineage: Lineage-Guided Enforcement for LLM Agent Memory, AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment[object Object][object Object]md
20260609_socratic_swe_self_evolving_coding_agents_trace_derived_skillspaper_analysisself_evolving_agents, memory_mechanism, skill_curation, reinforce_learning, reasoningsjtu, swebench_verified, terminalbench_2, grpo, skillos100Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skillsskillos_learning_skill_curation_self_evolving_agents, skillopt_executive_strategy_self_evolving_agent_skills, aflow_automating_agentic_workflow_generation, scaling_harness_agentic_ai, meme_multi_entity_evolving_memory_evaluation[object Object][object Object]md
20260627_joint_learning_experiential_rules_policies_llm_agentspaper_analysismemory_mechanism, reinforce_learning, self_evolving_agents, reasoning, agent_architecturealfworld, grpo, sun_yat_sen_university100Joint Learning of Experiential Rules and Policies for Large Language Model Agentsunlocking_working_memory_latent_reasoning, memory_os_of_ai_agent, memskill_learning_evolving_memory_skills[object Object][object Object]md
20260629_memory_r1_enhancing_llm_agents_manage_utilize_memories_rlpaper_analysisagent_architecture, memory_mechanism, reinforce_learning, llm, long_term_memorylmu_munich, mem0, locomo_bench, grpo100Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learningmemgpt_towards_llms_as_operating_systems, muse_autoskill_self_evolving_skill_memory, gems_agent_native_multimodal_generation_memory_skills, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications[object Object][object Object]md
20260701_severa_verified_synthesis_self_evolving_agentspaper_analysisagent_architecture, self_evolving_agents, reinforce_learning, reasoning, agent_securityuniversity_of_illinois_urbana_champaign, grpo, tau_squared_bench100SEVerA: Verified Synthesis of Self-Evolving Agentsskillos_learning_skill_curation_self_evolving_agents, muse_autoskill_self_evolving_skill_memory, trace_unified_rollout_budget_agentic_rl, amr_sd_token_level_credit_assignment, agent_memory_characterization_system_implications[object Object][object Object]md
20260706_demopsd_disagreement_modulated_policy_self_distillationpaper_analysisreinforce_learning, reasoning, llm, self_evolving_agentsgrpo, sciknoweval, kl_distillation, gpqa100DemoPSD: Disagreement-Modulated Policy Self-Distillationrl_long_horizon_reasoning, vector_policy_optimization[object Object][object Object]md
Powered by Forestry.md