Graduate-Level Google-Proof Q\u0026A benchmark for evaluating reasoning and scientific knowledge in LLMs
| Name | Categories | Topics | References | Credibility | CreateDate | UpdateDate | Title | RelatedNotes | NavigationOrder | Eleventy | TemplateEngineOverride | Created |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 20260706_demopsd_disagreement_modulated_policy_self_distillation | paper_analysis | reinforce_learning, reasoning, llm, self_evolving_agents | grpo, sciknoweval, kl_distillation, gpqa | 100 | DemoPSD: Disagreement-Modulated Policy Self-Distillation | rl_long_horizon_reasoning, vector_policy_optimization | [object Object] | [object Object] | md |