Evaluation methods, benchmarks, and assessment frameworks for AI systems

表格 14 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleRelatedNotesNavigationOrderEleventyTemplateEngineOverrideCreated
20260512_meme_multi_entity_evolving_memory_evaluationpaper_analysismemory_mechanism, long_term_memory, agent_architecture, llm, reasoning, evaluationkaist_ai, tuebingen_ai_center, naver_ai_lab100MEME: Multi-entity & Evolving Memory Evaluationmemory_os_of_ai_agent, memgpt_towards_llms_as_operating_systems, disentangling_memory_reasoning_llm[object Object][object Object]md
20260520_skillgenbench_benchmarking_skill_generation_llm_agentspaper_analysisagent_architecture, evaluation, skill_curation, llm, reasoningquanta_alpha, sjtu, pku, nus, xjtu, tsinghua_university, sufe, ntu, ucas100SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agentsontology_skill_analysis[object Object][object Object]md
20260529_agent_lifespan_engineering_for_deployed_systemspaper_analysismemory_mechanism, long_term_memory, agent_architecture, evaluation, self_evolving_agentsut_austin100Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed SystemsYour Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems[object Object][object Object]md
20260529_stop_comparing_llm_agents_without_disclosing_harnesspaper_analysisagent_architecture, evaluation, llm, multi_agent_systems, reasoningtulane_university, rutgers_university, virginia_tech100Stop Comparing LLM Agents Without Disclosing the HarnessStop Comparing LLM Agents Without Disclosing the Harness[object Object][object Object]md
20260601_locally_coherent_globally_incoherent_multi_component_llm_agentspaper_analysismulti_agent_systems, reasoning, agent_architecture, evaluationprinceton_university, paleka100Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agentsscaling_large_language_model_multi_agent_collaboration, gasim_graph_accelerated_hybrid_social_simulation, harnessing_agentic_evolution, scaling_harness_agentic_ai[object Object][object Object]md
20260603_agentcl_continual_learning_language_agentspaper_analysiscontinual_learning, memory_mechanism, agent_architecture, evaluation, self_evolving_agentsagentcl, ohio_state_university100AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agentslearning_fast_slow_llms_adapt_continually, meme_multi_entity_evolving_memory_evaluation, muse_autoskill_self_evolving_skill_memory[object Object][object Object]md
20260604_agent_memory_characterization_system_implicationspaper_analysismemory_mechanism, agent_architecture, long_term_memory, llm, evaluationstanford_university, ku_leuven, memory_agent_bench, memgpt, graphrag95Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloadsmemgpt_towards_llms_as_operating_systems, memlineage_lineage_guided_llm_agent_memory, calmem_dual_memory_conversational_ai, unlocking_working_memory_latent_reasoning, aip_graph_representation_for_learning_and_governing_agent_skills[object Object][object Object]md
20260622_how_transparent_is_diffusiongemmapaper_analysisllm, reasoning, agent_security, evaluation, interpretabilitygoogle_brain100How Transparent is DiffusionGemma?detecting_hallucinations_semantic_entropy[object Object][object Object]md
20260624_evoarena_tracking_memory_evolutionpaper_analysismemory_mechanism, agent_architecture, llm, evaluation, embodied_ainus, ntu, terminalbench_2100EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic EnvironmentsSelf-Harness: Harnesses That Improve Themselves[object Object][object Object]md
20260624_world_models_in_pieces_structural_certificationpaper_analysisreinforce_learning, reasoning, agent_architecture, evaluation, foundation_agentscuhk_shenzhen100World Models in Pieces: Structural Certification for General Agentsrl_long_horizon_reasoning_llm_expressiveness, agent_memory_characterization_system_implications[object Object][object Object]md
20260705_llm_agents_social_structure_latent_objective_emergencepaper_analysismulti_agent_systems, social_simulation, agent_architecture, reasoning, llm, evaluationcmu, dual_channel_debate, llm_agora100What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debatesself_harness, contagion_networks_evaluator_bias_multi_agent_llm[object Object][object Object]md
20260707_evopolicygym_evaluating_autonomous_policy_evolutionpaper_analysisembodied_ai, reinforce_learning, self_evolving_agents, evaluation, agent_architecturecuhk_shenzhen, sjtu, tsinghua_university, ustc100EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environmentsdemopsd_disagreement_modulated_policy_self_distillation, llm_agents_social_structure_latent_objective_emergence, recontext_recursive_evidence_replay, evolvenav_proactive_preflection_self_evolving_memory[object Object][object Object]md
20260708_adacurl_adaptive_curriculum_rlpaper_analysisreinforce_learning, llm, reasoning, agent_architecture, evaluationamap100AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical Revisiting[object Object][object Object]md
20260714_rubrics_as_rewards_reinforcement_learning_beyond_verifiable_domainspaper_analysisreinforce_learning, reward_modeling, llm, reasoning, evaluationscale_ai100Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains[object Object][object Object]md
Powered by Forestry.md