Graduate-Level Google-Proof Q\u0026A benchmark for evaluating reasoning and scientific knowledge in LLMs

表格 1 results
NameCategoriesTopicsReferencesCredibilityCreateDateUpdateDateTitleRelatedNotesNavigationOrderEleventyTemplateEngineOverrideCreated
20260706_demopsd_disagreement_modulated_policy_self_distillationpaper_analysisreinforce_learning, reasoning, llm, self_evolving_agentsgrpo, sciknoweval, kl_distillation, gpqa100DemoPSD: Disagreement-Modulated Policy Self-Distillationrl_long_horizon_reasoning, vector_policy_optimization[object Object][object Object]md
Powered by Forestry.md