Interactive science experiment benchmark for evaluating long-term memory, multi-step planning, and scientific reasoning in LLM agents.

表格 1 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleRelatedNotesNavigationOrderEleventyTemplateEngineOverrideCreated
20260717_experience_memory_graph_one_shot_error_correction_for_agentspaper_analysismemory_mechanism, agent_architecture, reasoning, long_term_memory, knowledge_graphalfworld, scienceworld, uestc100Experience Memory Graph: One-Shot Error Correction for Agentsmemlineage_lineage_guided_llm_agent_memory, shared_selective_persistent_memory_agentic_llm, agent_memory_and_reasoning_frontier_survey_2026[object Object][object Object]md
Powered by Forestry.md