← 返回 Avalaches

9月5日,人工智能在综合预测领域取得历史性突破:AI系统首次夺得Metaculus杯季度预测大赛冠军,并包揽了第二名和第五名,人类预测者仅获得第三和第四名。数百名参赛者围绕地缘政治、大宗商品价格及疫情传播等具体问题展开数月推演,系统依据预测偏离实际结果的距离与时间跨度综合计分。除5,000美元的比赛奖金外,算法在真实预测市场展现出更为惊人的盈利能力:自6月以来,FutureSearch利用其AI在Kalshi预测市场中将10万美元初始组合增值了6%;而另一家未进前五的团队开发者更是创造了Kalshi历史第六高的回报率,在七个月内将35美元暴赚至194万美元。

现代预测科学正经历从统计分析向多模态大语言模型(LLM)的范式演进。宾夕法尼亚大学2015年的经典研究表明,顶尖的人类“超级预测者”对未来300天事件的预判精确度等同于普通人对未来60天的判断;而预测研究所(FRI)今年7月的最新评估显示,AI系统在动态演化的预测题库中已与超级预测者达到同等水平。顶尖的预测AI并非由巨额资金垄断,独立开发者杰弗里·梁(Jeffrey Liang)仅耗时不足150小时、花费数千美元算力与数据打造的算法模型,便击败了累计融资超1,500万美元的四家专业初创企业,并在奖金达50,000美元的纯AI预测大赛中领跑。

相比于耗资逾10,000美元且耗时一周的人类专家研判,FutureSearch等AI工具仅需数分钟和数美元即可生成详尽的概率论证。AI系统能够吸收并交叉辩论不同模型的逻辑链,利用付费数据库回溯清洗历史指令偏误,展现出无可匹敌的归纳综合效率。尽管人类研判在跨越数年的宏观定性预见上仍保有微妙优势,但AI的极速演进正在重塑决策咨询生态,甚至有望厘清长期预测中混沌固有的不可知边界与技术性认知的客观极限。

Artificial intelligence now beats some of the best human forecasters image

On September 5th artificial intelligence breached another frontier of human cognitive superiority as algorithms dominated the seasonal Metaculus Cup forecasting competition, capturing first, second, and fifth places while relegating human participants to third and fourth. Across hundreds of entries evaluating complex real-world variables—including data-centre regulation, disease transmission, and Brent crude movements—participants were scored on the continuous calibration of their probabilities against realized outcomes. Beyond the contest’s ,000 cash pool, algorithmic models demonstrated formidable real-world financial returns: platform FutureSearch expanded its initial ,000 Kalshi prediction-market portfolio by 6% since June, while an individual developer achieved the sixth-highest return in Kalshi’s history by compounding into .94m over seven months.

This quantitative shift marks the convergence of general large language models (LLMs) with high-stakes forecasting methodologies. Foundational 2015 research in Perspectives on Psychological Science demonstrated that elite human “superforecasters” could predict outcomes 300 days out with accuracy matching regular forecasters assessing a 60-day horizon; by July, independent evaluations by the Forecasting Research Institute confirmed that advanced AI architectures have achieved full parity with these human superforecasters. Crucially, competitive edge does not rest solely on capital scale: independent developer Jeffrey Liang utilized under 150 development hours and several thousand dollars in computing resources to outperform four venture-backed startups possessing over m in aggregate funding, subsequently taking the lead in a ,000 AI-exclusive tournament.

Forecasting economics are undergoing radical democratization: whereas commissioned human superforecaster assessments routinely cost upwards of ,000 and require a week of analysis, emerging automated platforms deliver multi-model probability breakdowns in ten minutes for nominal fees. Systemic capabilities—such as cross-model dialectical debate, proprietary data integration, and total memory resets to evaluate counterfactual prompts against past events—enable iterative feedback loops impossible for human minds. While human qualitative judgment still retains an advantage over multi-year horizons where empirical precedent is thin, machines will increasingly synthesise sprawling global data inputs, simultaneously illuminating where true atmospheric chaos begins and what parts of the future remain fundamentally knowable.

Source: Artificial intelligence now beats some of the best human forecasters

Subtitle: Crystal balls give way to LLMs

Dateline: Sep 17th 2026\n


2026-09-19 (Saturday) · 8e7b9fbed0c4ebbc1f8d644c5f158d2f2dc5a3bd

Attachments