Dynamic Advantage Policy Optimization, a reinforcement learning algorithm for LLMs

表格 2 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleRelatedNotesNavigationOrderEleventyTemplateEngineOverrideCreated
20260422_multimodal_agent_long_term_memorypaper_analysisagent_architecture, multimodal, memory_mechanism, long_term_memory, video_understandingm3_bench, bytedance_seed, dapo100Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryHippoRAG: Neurobiologically Inspired Long-Term Memory[object Object][object Object]md
20260520_amr_sd_token_level_credit_assignmentpaper_analysisreinforce_learning, reasoning, llm, reward_modeling, self_evolving_agentsgrpo, dapo, meituan, sciknoweval100AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignmentrl_long_horizon_reasoning_llm_expressiveness, skillos_learning_skill_curation_self_evolving_agents[object Object][object Object]md
Powered by Forestry.md