BFCL-v3 (Berkeley Function Calling Leaderboard v3) benchmark for function calling capabilities
| Name | Categories | Topics | References | Credibility | CreateDate | UpdateDate | Title | NavigationOrder | Eleventy | TemplateEngineOverride | Created | RelatedNotes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 20260423_agentevolver_self_evolving_agent | paper_analysis | agent_architecture, reinforce_learning, llm, reasoning | tongyi_lab, appworld, bfcl_v3 | 100 | AgentEvolver: Towards Efficient Self-Evolving Agent System | [object Object] | [object Object] | md | ||||
| 20260611_trace_unified_rollout_budget_allocation_agentic_rl | paper_analysis | reinforce_learning, reasoning, agent_architecture, test_time_scaling, llm | tsinghua_university, hotpotqa, bfcl_v3, monte_carlo_tree_search | 100 | TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning | [object Object] | [object Object] | md | rl_long_horizon_reasoning_llm_expressiveness, agent_memory_and_reasoning_frontier_survey_2026 |