磨砺预测技能:远见是一种可培养的可衡量技能

2015 · report · 原文约 9432 词
译文与英文原文逐段对齐可在本页展开英文,也可打开发布者原址核对上下文。
打开来源正文

全球金融策略 www.credit-suisse.com

GLOBAL FINANCIAL STRATEGIES www.credit-suisse.com

提升你的预测技能 远见是一项可衡量的技能,你可以培养它 2015 年 9 月 28 日

Sharpening Your Forecasting Skills Foresight Is a Measurable Skill That You Can Cultivate September 28, 2015

Authors

Authors

迈克尔·莫布森 [email protected]

Michael J. Mauboussin [email protected]

丹·卡拉汉,特许金融分析师(CFA),[email protected]

Dan Callahan, CFA [email protected]

“信念是需要检验的假设,不是需要保护的宝藏。”

“Beliefs are hypotheses to be tested, not treasures to be protected.”

Philip E. Tetlock 与丹·加德纳

Philip E. Tetlock and Dan Gardner1

菲利普·泰特洛克对数百名专家在二十年间做出的数千次预测进行的研究发现,这些预测的平均水平“并不比瞎猜好多少”。这就是坏消息。

Philip Tetlock’s study of hundreds of experts making thousands of predictions over two decades found that the average prediction was “little better than guessing.” That’s the bad news.

泰特洛克与他的同事们参加了一场由美国情报界赞助的预测竞赛。这项工作识别出了“超级预测者”——那些总能做出卓越预测的人。这是个好消息。

Tetlock, along with his colleagues, participated in a forecasting tournament sponsored by the U.S. intelligence community. That work identified “superforecasters,” people who consistently make superior predictions. That’s the good news.

超级预测者们的关键在于他们的思维方式。他们积极保持思想开放、具备知识上的谦逊、精通数据、善于反思地更新自己的判断,并且工作极为勤奋。

The key to superforecasters is how they think. They are actively open-minded, intellectually humble, numerate, thoughtful updaters, and hard working.

超级预测者在团队中工作时能取得更好的成果。但由于团队协作有利有弊,培训就显得至关重要。

Superforecasters achieve better results when they are part of a team. But since there are pros and cons to working in teams, training is essential.

教授减少预测偏差的方法能改善结果。

Instruction in methods to reduce bias in forecasts improves outcomes.

培训与执行之间必须有紧密的联系。

There must be a close link between training and implementation.

最优秀的领导者明白,恰当甚至大胆的行动,需要有好的思考作为前提。

The best leaders recognize that proper, even bold, action requires good thinking.

导言:坏消息与好消息

Introduction: The Bad News and the Good News

如果你有机会学习如何将自己预测的准确性——以预测与实际结果的差距来衡量——提升 60%,你愿意吗?感兴趣吗?菲利普·泰特洛克与丹·加德纳合著的《超预测:预测的艺术与科学》一书,展示了少数“超级预测者”是如何达到这种技能水平的。如果你从事预测行业——既然你在读这段话,那你很可能就是——你现在就该抽出一点时间去买这本书。你会发现,这是一本既基于科学又极具实用价值的难得之作。

What if you had the opportunity to learn how to improve the quality of your forecasts, measured as the distance between forecasts and outcomes, by 60 percent? Interested? Superforecasting: The Art and Science of Prediction by Philip Tetlock and Dan Gardner is a book that shows how a small number of “superforecasters” achieved that level of skill. If you are in the forecasting business—which is likely if you’re reading this—you should take a moment to buy it now. You’ll find that it’s a rare book that is both grounded in science and highly practical.

菲尔·泰特洛克是宾夕法尼亚大学的心理学与政治学教授,他花了数十年时间研究专家们的预测。具体来说,他吸引 284 位专家在 2004 年结束的 21 年间,对政治、社会和经济结果做出了超过 2.7 万个预测。这段时期涵盖了六次总统选举和三场战争。这些预测者拥有极其亮眼的资历,包括十多年的相关工作经验以及大量高等学位——几乎所有人都接受过研究生训练,一半人拥有博士学位。

Phil Tetlock is a professor of psychology and political science at the University of Pennsylvania who has spent decades studying the predictions of experts. Specifically, he enticed 284 experts to make more than 27,000 predictions on political, social, and economic outcomes over a 21-year span ended in 2004. The period included six presidential elections and three wars. These forecasters had crack credentials, including more than a dozen years of relevant work experience and lots of advanced degrees—nearly all had postgraduate training and half had PhDs.

接着,泰特洛克做了一件非常不寻常的事:他持续记录了专家们的预测。结果汇总在他那本《专家的政治判断》中,令人沮丧。² 普通专家的预测“比瞎猜好不了多少”,这算是委婉说法,直白点就是“准确度大致相当于一只扔飞镖的黑猩猩”。面对自己预测无用的证据,这些专家做出了和我们所有人一样的事:竖起了心理防御盾牌。他们说自己差点就对了,或者自己的预测分量太重以至于影响了结果,再或者预测本身没错,只是时机掐错了。总体来看,泰特洛克的研究结果,为那些质疑专家价值的人提供了极具杀伤力的弹药。

Tetlock then did something very unusual. He kept track of their predictions. The results, summarized in his book Expert Political Judgment, were not encouraging.2 The predictions of the average expert were “little better than guessing,” which is a polite way to say that “they were roughly as accurate as a dart-throwing chimpanzee.” When confronted with the evidence of their futility, the experts did what the rest of us do: they put up their psychological defense shields. They noted that they almost called it right, or that their prediction carried so much weight that it affected the outcome, or that they were correct about the prediction but simply off on timing. Overall, Tetlock’s results provide lethal ammunition for those who debunk the value of experts.

在“专家无效性”这个头条结论之下,还有一些更微妙的发现。其一是名气与准确率呈反向关联。那些声名显赫的专家预测记录堪称最差,却展现出“讲述引人入胜故事的本领”。要想出名,就得会讲“紧凑、简单、清晰、能抓住并维系听众注意力的故事”。这些权威人士常常犯错,却从不自我怀疑。

Below the headline of expert ineffectiveness were some more subtle findings. One was an inverse correlation between fame and accuracy. While famous experts had among the worst records of prediction, they demonstrated “skill at telling a compelling story.” To gain fame it helps to tell “tight, simple, clear stories that grab and hold audiences.” These pundits are often wrong but never in doubt.

另一个与前述相关的发现是:决定预测质量的,更多在于专家如何思考,而非他们想什么。泰特洛克依据哲学家以赛亚·伯林关于思维风格的名篇,将专家分为狐狸型和刺猬型。狐狸型略知万物,刺猬型独精一事。狐狸型表现优于投飞镖的猩猩,而刺猬型则更差。

Another result, which is related to the first, was that what mattered in the quality of predictions was less what the expert thought and more how he or she thought. Tetlock categorized his experts as foxes or hedgehogs based on a famous essay on thinking styles by the philosopher Isaiah Berlin. Foxes know a little about a lot of things, and hedgehogs know one big thing. Foxes did better than the dart-throwing chimp, and hedgehogs did worse.

不难看出这些发现之间的关联。经济、社会和政治领域中的大多数热门话题,都无法被紧密、简单且清晰的故事所概括。但设想一下,如果你是一档政治类电视节目的制作人,你希望邀请哪位嘉宾上镜——是那位模棱两可、总是说“另一方面”的嘉宾,还是那位自信满满、讲述一个干脆利落又富有争议故事的嘉宾?这并非一个艰难的决定,正因如此,许多“刺猬型”人物既声名远扬,又预测能力糟糕。

It’s not hard to see the link between these findings. Most topics of interest in the economic, social, and political realms defy tight, simple, and clear stories. But imagine you are the producer of a television show that covers politics. Who do you want to put on the air, the equivocal guest who constantly says “on the other hand,” or the one who confidently tells a crisp and controversial story? It’s not a hard decision, which is why many hedgehogs are both famous and poor predictors.

尽管《专家政治判断》的结论颇为精微,但总体而言,这对那些评论员来说是个坏消息。

While the conclusions of Expert Political Judgment were nuanced, they were on balance bad news for pundits.

尽管有些人从他的研究结果中读出了极端结论,但泰特洛克从不认为预测毫无用处——狐狸比所有专家的平均水平预测得更准,这本身就是一个有力线索,说明预见力或许是一种可以被识别和培养的真实技能。泰特洛克把自己定位为“乐观的怀疑论者”。

Despite how some read his results, Tetlock never believed in the extreme point of view that forecasts are useless. That foxes were better forecasters than the average of all experts provided a strong clue that foresight might be a real skill that could be identified and cultivated. Tetlock marked himself as an “optimistic skeptic.”

《专家的政治判断》是一项优秀的学术研究成果,但它是用——嗯——学术腔写成的。在《超预测》一书中,泰特洛克与记者丹·加德纳合作,后者曾写过一本关于预测失败的著作。成果是一本易读又扎实的研究著作。

Expert Political Judgment is excellent scholarly research but is written in, well, scholarly prose. In Superforecasting, Tetlock collaborates with Dan Gardner, a journalist and author of a book about the failure of prediction. The result is great research that is easy to read.

自然,泰特洛克并不是唯一一个对学习如何做出有效预测感兴趣的人。美国情报界也迫切希望提高预测质量,尤其是在未能预见 2001 年 9 月 11 日恐怖袭击事件,以及 2003 年高估伊拉克存在大规模杀伤性武器可能性之后。情报界下属的机构——情报高级研究计划局(IARPA)——应运而生,致力于开展高风险研究,以探索如何提升美国情报水平。IARPA 决定举办一场预测锦标赛,想看看是否存在某种方法能提高预测的精准度。

Naturally, Tetlock is not the only one interested in learning how to make effective forecasts. The United States intelligence community was also keen to improve the quality of predictions, especially in the wake of the failure to anticipate the terrorist acts on September 11, 2001 and the overestimation of the probability of the existence of weapons of mass destruction in Iraq in 2003. An agency within the community, Intelligence Advanced Research Projects Activity (IARPA), was assembled to pursue high-risk research into how to improve American intelligence. IARPA decided to create a forecasting tournament to see if there might be a way to sharpen forecasts.

泰特洛克与几位同事共同发起了“良好判断项目”(GJP),这是参与准确回答问题的五支科学团队之一。这些团队可以使用任何想得到的方法来得出最佳答案。从 2011 年 9 月开始,IARPA 提出了近 500 个关于各种政治与经济结果的问题。这场竞赛在随后的四年里累计收集了超过一百万份个人预测。值得注意的是,IARPA 竞赛中问题的时间跨度通常为一个月至一年,比泰特洛克研究专家时所常见的三到五年要短。

Tetlock and some colleagues launched the Good Judgment Project (GJP), one of five scientific teams that would compete to answer questions accurately. The teams could use whatever approaches they wanted to generate the best possible answers. Starting in September 2011, IARPA asked nearly 500 questions about various political and economic outcomes. The tournament garnered more than one million individual forecasts in the following four years. It is important to note that the time frames for the questions in the IARPA tournament, generally one month to one year, were shorter than the three to five years that were common in Tetlock’s study of experts.

好消息来了:GJP 项目第一年的成绩比对照组高出 60%。第二年的结果更出色,几乎以 80% 的优势碾压对照组。事实上,GJP 表现之优异,使得 IARPA 直接解散了其他团队。

Now the good news: the GJP results beat the control group by 60 percent in year one. Results in year two were even better, trouncing the control group by almost 80 percent. In fact, the GJP did so well that IARPA dropped the other teams.

在比赛第一年的 2,800 名 GJP 志愿者中,表现最顶尖的 2% 被称为“超级预测者”。据《华盛顿邮报》一位编辑的说法,从预测能力来看,这些超级预测者的表现比情报界——那些有机会接触机密数据的人——的平均水平还要高出大约 30%。

Of the 2,800 GJP volunteers in the first year of the tournament, the top 2 percent were called “superforecasters.” To give you some sense of their acuity, the superforecasters performed about 30 percent better than the average for the intelligence community—people who had access to classified data—according to an editor at the Washington Post.3

受到 GJP 研究成果的鼓舞,泰特洛克得出了几个结论。第一个是,预见力是一种真实且可测量的技能。衡量技能的指标之一是持续性。持续性高意味着随着时间的推移,你始终表现稳定,而不是昙花一现。大约 70% 的超级预测者能从一年到下一年一直保持在精英行列,这个比例远超随机概率。

Encouraged by the GJP’s results, Tetlock came to a couple conclusions. The first is that foresight is a real and measurable skill. One test of skill is persistence. High persistence means that you do consistently well over time and are not a one-hit wonder. About 70 percent of superforecasters remain in those elite ranks from one year to the next, vastly more than what chance would dictate.

第二点是,远见“是特定思维方式、信息收集方式和信念更新方式的产物”。重要的是,成为超级预测家的核心要素是可以通过学习和培养获得的。GJP 的美妙之处在于它以科学严谨的方式开展,这使得研究人员能够提炼出成功的要素。我们将在本报告中探讨这些要素。

The second is that foresight “is the product of particular ways of thinking, of gathering information, of updating beliefs.” Importantly, the essential ingredients of being a superforecaster can be learned and cultivated. The beauty of the GJP is that it was carried out with scientific rigor, which allowed the researchers to distill the elements of success. We explore these elements in this report.

尽管大多数人能够改善自己的思维和预测水平,但基于泰特洛克和加德纳所谓的“知识幻觉”,人们始终对改变抱有抵触。直觉就是其中之一。直觉是一种模式识别,在充满“有效线索”的环境下能够发挥作用。但直觉在不稳定或非线性环境中极不可靠。过度依赖直觉会导致糟糕的决策。

Even though most people can improve their thinking and forecasting, there has always been resistance to change based on what Tetlock and Gardner call “illusions of knowledge.” Intuition is one example. Intuition is a form of pattern recognition that works in settings with lots of “valid cues.”4 But intuition is notoriously unreliable in unstable or nonlinear environments. An overreliance on intuition leads to poor decisions.

另一个例子是缺乏自我反思。这在一定程度上是由我们大脑中一个模块所驱动的,它试图快速闭合因果循环。当我们向你展示一个结果时,你的大脑会迅速为其想出一个解释。正如泰特洛克和加德纳所写,“我们太快地从困惑和不确定跳到一个清晰且自信的结论,中间没有花任何时间。”这与著名心理学家丹尼尔·卡尼曼所称的“快速思考”概念有关。

Another case is insufficient self-reflection. This is in part prompted by a module in our brain that seeks to rapidly close cause-and-effect loops. We show you an outcome and your mind quickly comes up with an explanation for it. As Tetlock and Gardner write, “we move too fast from confusion and uncertainty to a clear and confident conclusion without spending any time in between.” This is related to the concept that Daniel Kahneman, an eminent psychologist, calls thinking fast.5

跟踪预测与实际结果,对专家的利益没什么好处。如果你是靠当权威评论员领高薪的,引入一套评分系统几乎没什么收益,反而风险极大。极端情况下,专家对自己的判断笃信不疑,甚至觉得压根没必要去衡量结果。

Keeping track of forecasts and the outcomes may not serve an expert’s interests. If you are paid well to be a pundit, introducing a scoring system offers little upside and lots of downside. In extreme cases, experts are so sure that they are correct that they see no need to measure outcomes at all.

泰特洛克和加德纳引用了公元二世纪罗马皇帝的御医盖伦的话。针对某种特定疗法,盖伦写道:“所有服用此药者都在短期内康复,唯独那些不见效的人,全都死了。由此可见,它只在不治之症上失败。”

Tetlock and Gardner share the words of Galen, the physician to Roman emperors, who practiced in the second century. Of a particular cure he wrote, “All who drink of this treatment recover in a short time, except those whom it does not help, who all die. It is obvious, therefore, that it fails only in incurable cases.”

幸运的是,近几个世纪以来,医学研究人员更严格地运用了科学方法,但要想发现一个过度自信且不受约束的专家,仍然很容易。

Fortunately, medical researchers have applied the scientific method more rigorously in recent centuries, but it’s still easy to spot an overconfident, and unchecked, expert.

那么,好的预测从何而来?泰特洛克及其同事发现了超级预测者成功的四个驱动因素:⁶

So what is the source of good forecasting? Tetlock and his colleagues found four drivers behind the success of the superforecasters:6

找到合适的人。通过对预测者的流体智力和主动开放性思维进行筛选,你能获得 10% 到 15% 的提升。

Find the right people. You get a 10-15 percent boost from screening forecasters on fluid intelligence and active open-mindedness.

管理互动。让预测者在团队中协作工作,或者在预测市场中竞争,可以获得 10% 到 20% 的提升。

Manage interaction. You get a 10-20 percent enhancement by allowing the forecasters to work collaboratively in teams or competitively in prediction markets.

有效训练。认知去偏误练习能让结果提升 10%。

Train effectively. Cognitive debiasing exercises lift results by 10 percent.

精英预测者应被赋予更高权重,或对估算进行极端化处理。若对更出色的预测者给予更大权重,并为弥补预测中的保守倾向而将预测结果推向更极端的区间,整体表现可提升 15% 至 30%。

Overweight elite forecasters or extremize estimates. Results improve by 15-30 percent if you give more weight to better forecasters and make forecasts more extreme to compensate for the conservatism of forecasts.

科学家们使用布赖尔评分(Brier score,附录中提供了更详细的计算方法)来衡量这些改进。布赖尔评分反映了预测与实际结果之间的差距。和高尔夫球得分一样,分值越低越好。计算布赖尔评分有几种方式,但常用标准范围是从 0 到 2.0。

The scientists measure these improvements using a Brier score (the appendix provides more detail on the calculation). A Brier score reflects the difference between a forecast and the outcome. Like golf scores, lower is better. There are a couple of ways to calculate Brier scores, but a common scale runs from zero to 2.0.

0 意味着预测完全准确,0.50 是随机预测,而 2.0 则表示预测完全错误。

Zero means that the forecast is spot on, 0.50 is a random forecast, and 2.0 means that the forecast is completely wrong.

依照这种评分方式,一位预测某结果发生概率为 55% 的人,若实际果然发生,其布赖尔得分为 0.405。随后他预测另一事件发生概率为 65%,若事件实际发生,布赖尔得分为 0.245,几乎改善了 40%。

By this scoring, a person who predicts a 55 percent probability of an outcome that happens receives a Brier score of 0.405. A subsequent forecast of a 65 percent probability of an event that occurs gets a Brier score of 0.245, nearly a 40 percent improvement.

找到合适的人、有效组建和管理团队、以及进行恰当培训,涉及大量细节。但泰洛克和加德纳给出了一个简洁公式,它是整个流程的核心:“预测、衡量、修正:这是通向更清晰洞察的最可靠之路。”

There are a lot of details in finding the right people, effectively building and managing teams, and proper training. But Tetlock and Gardner offer a simple formula that is at the core of the whole process: “Forecast, measure, and revise: it is the surest path to seeing better.”

我们如何都能变得更像超级预测者?回答这个问题,我们需要讨论超级预测者的画像、像超级预测者一样思考所需的工具、给领导者的经验教训,以及仍然存在的合理疑虑。

How can we all become more like superforecasters? To answer that question, we discuss the profile of a superforecaster, the tools you will need to think like a superforecaster, the lessons for leadership, and the valid doubts that remain.

找到合适的人

Find the Right People

整本书中,泰洛克和加德纳提供了各种碎片,让你能拼凑出超级预测者的画像。因为良好判断力项目(GJP)的研究人员花时间让预测者接受了一系列心理测试,他们得以审视超级预测者的个性。更进一步,这些数据

Throughout the book, Tetlock and Gardner provide the pieces that allow you to construct the profile of a superforecaster. Because the GJP researchers took the time to run the forecasters through a battery of psychological tests, they were able to examine the personalities of the superforecasters. Further, the data

允许研究人员避免先看到成功、再事后寻找共同特征的错误。典型超级预测者的画像有四个要素:

allow the researchers to avoid the error of first observing success and then attempting to find common attributes after the fact.7 The portrait of a modal superforecaster has four elements:

哲学观。 超级预测者往往能与怀疑感安然共处。科学家有时会感觉自己掌握了真相。优秀的思考者也会有同样感受。“但他们知道,必须把这种感觉放在一边,代之以精确定量的怀疑,”泰洛克和加德纳写道,“——这种怀疑可以通过更好的研究、更好的证据来减少(尽管永远无法减为零)。”

Philosophical Outlook. Superforecasters tend to be comfortable with a sense of doubt. Scientists sometimes sense that they know the truth. Good thinkers can feel the same way. “But they know they must set that feeling aside and replace it with finely measured degrees of doubt,” write Tetlock and Gardner, “— doubt that can be reduced (although never to zero) by better evidence from better studies.”

回想一下,我们的大脑热衷于归因。我们希望案子结案。但正如丹尼尔·卡尼曼所说,“接受不确定性是明智的,但高置信度的声明主要告诉你,这个人在脑子里编了一个连贯的故事,并不一定意味着这个故事是真的。”

Recall that our minds are keen to assign causality. We want the case to be closed. But as Daniel Kahneman says, “It is wise to take admissions of uncertainty seriously, but declarations of high confidence mainly tell you that an individual has constructed a coherent story in his mind, not necessarily that the story is true.”8

超级预测者也很谦逊,但这不是指感觉自身不配。相反,他们的谦逊来自于认识到现实极度复杂。事实上,完全可以在高度评价自己的同时,保持智识上的谦逊。泰洛克和加德纳指出:“智识谦逊促使你深思熟虑,这是良好判断所必需的;而对自己能力的信心则激发果断行动。”

Superforecasters are also humble, but not in the sense of feeling unworthy. Rather, their humility comes from the recognition that reality is profoundly complex. Indeed, it is possible to think highly of yourself and to be intellectually humble at the same time. Tetlock and Gardner note that, “Intellectual humility compels the careful reflection necessary for good judgment; confidence in one’s abilities inspires determined action.”

把结果归因于命运很常见,也常常令人心安。超级预测者不太相信命运。

It is common, and often soothing, to attribute outcomes to fate. Superforecasters aren’t big believers in fate.

在一个从 1 到 9 的“命运分数”上,1 代表完全否定命运,9 代表完全相信命运,普通美国成年人平均分落在中间。宾夕法尼亚大学学生的平均分略低一点,普通预测者低于学生,而超级预测者在这几组人中分数最低。超级预测者不认为发生的事就注定要发生。

On a one to nine “fate score,” where one is a total rejection of fate and nine is complete belief in it, the average adult American falls near the middle. The mean score for a student at the University of Pennsylvania is a little lower, the regular forecasters are below that, and the superforecasters are the lowest of these groups. Superforecasters don’t think that what happened had to happen.

能力和思维方式。 第一点是,超级预测者并非天才。研究人员测试了所有 GJP 志愿者的流体智力和晶体智力。流体智力是逻辑思考和解决新问题的能力,不依赖积累的知识。晶体智力顾名思义,是你掌握的技能、事实和智慧,以及你在需要时运用它们的能力。

Ability and Thinking Style. The first point is that superforecasters are not geniuses. The researchers tested the fluid and crystallized intelligence of all the GJP volunteers. Fluid intelligence is the ability to think logically and to solve novel problems. It doesn’t rely on accumulated knowledge. Crystallized intelligence is exactly what it sounds like: your collection of skills, facts, and wisdom, and your ability to use them when you need to.

参与 GJP 的人并非一个有效的人口样本——这些人是主动举手,为了换取 250 美元亚马逊礼品卡而做大量预测的人。普通预测者在智力测试中得分高于约 70% 的人口。这大致相当于平均智商在 108-110 之间,而人口平均值为 100。超级预测者得分高于约 80% 的人口,平均智商范围在 112-114 之间。

Those who participated in the GJP were not a valid sample of the population—these are people who raised their hand to make lots of forecasts in return for a $250 gift certificate from Amazon.com. The regular forecasters scored higher than about 70 percent of the population on intelligence tests. That translates roughly into an average intelligence quotient (IQ) of 108-110 where the average of the population is 100. The superforecasters scored higher than about 80 percent of the population, or an average IQ range of 112-114.

总体人口与普通预测者之间的差距,远大于普通预测者与超级预测者之间的差距。

There is a much bigger gap between the overall population and regular forecasters than there is between those forecasters and the superforecasters.

多伦多大学应用心理学与人类发展荣誉教授基思·斯塔诺维奇区分了智商和他所称的“理商”(Rationality Quotient, RQ)。两者之间的相关系数相对较低,在 0.20 到 0.35 之间。高理商的人展现出适应性行为、高效的行为调节、合理的目标优先级设定、反思性以及正确对待证据的能力。这些特质与超级预测者高度吻合。

Keith Stanovich, professor emeritus of applied psychology and human development at the University of Toronto, distinguishes between IQ and what he calls “RQ,” or rationality quotient.9 The correlation coefficient between the two is a relatively low .20 to .35. Those with high RQ’s exhibit adaptive behavioral acts, efficient behavioral regulation, sensible goal prioritization, reflectivity, and the proper treatment of evidence. These qualities are very consistent with those of the superforecasters.

宾夕法尼亚大学心理学教授、泰洛克的同事乔纳森·巴伦创造了一个术语:“主动开放性思维”。那些主动开放思维的人会寻找与自己不同的观点,并仔细审视它们。泰洛克和加德纳表示,如果他们必须把超级预测术减缩成一条汽车保险杠贴纸,那应该是:“信念是需要检验的假设,而非需要守护的宝藏。”

Jonathan Baron, a professor of psychology at the University of Pennsylvania and a colleague of Tetlock’s, coined the term “active open-mindedness.”10 Those who are actively open-minded seek views that are different than their own and consider them carefully. Tetlock and Gardner suggest that if they had to reduce superforecasting to a bumper sticker, it would read, “Beliefs are hypotheses to be tested, not treasures to be guarded.”

“大五人格”是接受度最广的人格测试之一。受试者接受五大性格特征的测试:经验开放性、尽责性、外向性、宜人性、神经质。测试显示,超级预测者在经验开放性上得分很高,这表明他们偏好认知多样性和智识好奇心。超级预测者对世界充满兴趣,是愿意探索的人。

The Big Five is one of the most widely-accepted personality tests. Subjects are tested for five personality traits: openness to experience, conscientiousness, extraversion, agreeableness, and neuroticism. The tests revealed that superforecasters score high in openness to experience, which suggests a preference for cognitive variety and intellectual curiosity. Superforecasters are interested in the world and are willing explorers.

超级预测者还会花时间思考自己的思考过程,并不断寻求改进。在协作时,超级预测者经常会在线上讨论中留下大量评论,这让他们能够重现自己的思维过程,并在可能时加以改进。及时准确的反馈是改进的关键要素。超级预测者拥抱反馈。

Superforecasters also spend time thinking about their own process and constantly seek to improve. When collaborating, the superforecasters often leave lots of comments in their online discussions, which allow them to recreate their thought processes and improve them when possible. Timely and accurate feedback is an essential element of improvement. Superforecasters embrace feedback.

超级预测者很少使用复杂的数学模型来做预测,但他们无一例外都拥有很高的数感。擅长数字是做出好预测的前提,但花哨的定量模型则不是。

The superforecasters rarely use sophisticated mathematical models to make their forecasts, but they are uniformly highly numerate. Comfort with numbers is a prerequisite for making good forecasts but fancy quantitative models are not.

预测方法。 超级预测者,类似于泰洛克对专家政治判断研究中的“狐狸”,在方法上倾向于务实。好的预测不是只用单一视角看待世界的所有方面,而是需要考虑多种观点。伯克希尔哈撒韦副董事长查理·芒格很好地用“思维模型”方法捕捉了这一概念。芒格说:“嗯,第一条规则是,你必须拥有多个模型——因为如果你只用一两个,人类心理的本性就会让你扭曲现实,使之符合你的模型,或者至少你会这么认为。”

Methods of Forecasting. Superforecasters, similar to the foxes in Tetlock’s study of expert political judgment, tend to be pragmatic in their methods. Rather than looking at all aspects of the world through a single lens, good forecasting requires considering multiple points of view. Charles Munger, vice chairman of Berkshire Hathaway, captures this concept well with the mental models approach. Says Munger, “Well, the first rule is that you’ve got to have multiple models—because if you just have one or two that you’re using, the nature of human psychology is such that you'll torture reality so that it fits your models, or at least you’ll think it does.”11

卡尼曼普及了大脑两个系统的概念。系统 1 快速、自动、难以训练。系统 2 缓慢、深思熟虑、有明确目的。这项研究的一个重要结论是,我们常常在我们本应调用慢系统时,却依赖我们的快系统。泰洛克和加德纳称快系统为“鼻尖视角”,因为它对我们每个人来说都是独一无二的。超级预测者清楚地知道何时需要调用系统 2。

Kahneman has popularized the notion of two systems of the mind. System 1 is fast, automatic, and difficult to train. System 2 is slow, deliberate, and purposeful. One of the important conclusions from this research is that we commonly rely on our fast system when we should recruit our slow system.12 Tetlock and Gardner call the fast system the “tip-of-your-nose perspective,” because it is unique to each of us. Superforecasters have a firm sense of when they need to engage System 2.

有充分证据表明,正确聚合不同观点可以提高预测的准确性。詹姆斯·索罗维基在他的畅销书《群体的智慧》中提供了大量例证。你可以通过收集不同个体(例如股市中的投资者)的观点,或者通过在你头脑中聚合多种观点,来从多样性中获益。

There is good evidence that the aggregation of diverse points of view, done correctly, improves the accuracy of forecasts. James Surowiecki provides ample illustrations of this idea in his bestselling book, The Wisdom of Crowds.13 You can gain from diversity by capturing the views of different individuals, for example, investors in a stock market, or by aggregating multiple views in your head.

泰洛克和加德纳使用了蜻蜓眼睛的比喻。每只眼睛有多达 3 万个独立的小眼,朝向略微不同的方向,为蜻蜓的大脑提供海量输入。结果是超凡的视觉敏锐度,使蜻蜓能够捕捉到小而快速移动的昆虫。

Tetlock and Gardner use the metaphor of a dragonfly’s eye. Each eye has up to 30,000 individual lenses aimed in slightly different directions that provide the dragonfly’s brain with massive input. The result is extraordinary visual acuity, allowing the dragonfly to nab small, fast-moving insects.

作者发现,超级预测者像蜻蜓一样,能够考虑并综合多种观点。他们还强调,“聚合并非我们的本能。”我们通常满足于自己的信念,认为没有必要去考虑替代性想法。有时答案是调查他人的观点,并请他们批评你的观点。另一些时候,你可以简单地在不同时间思考同一个主题,在自己的头脑中创造一个“群体”。无论你如何做到,吸纳他人的观点都是有价值的。

The authors find that superforecasters, similar to the dragonfly, are able to consider and synthesize multiple points of view. They also emphasize that “aggregation doesn’t come to us naturally.” We are generally content with our own beliefs and see no reason to entertain alternative thoughts. Sometimes the answer is to survey the views of others and ask them to criticize your view. Other times you can simply think about the same topic at different times and create a crowd within your head. No matter how you get there, taking in the views of others is valuable.

优秀的预测者用概率思考,但这绝非易事。泰洛克和加德纳提出,我们出厂时,概率刻度盘上只有三档:一定会发生;一定不会发生;可能会发生。他们指出,这对我们的祖先来说效果很好。“那是一只狮子吗?是 = 快跑!可能会 = 保持警惕!否 = 放松。”

Good forecasters think in probabilities, but there’s nothing easy about that. Tetlock and Gardner propose that we come out of the factory with three settings on our dial of probability: it’s going to happen; it’s not going to happen; and maybe. They suggest this worked fine for our ancestors. “Is that a lion? YES = run! MAYBE = stay alert! NO = relax.”

预测者常用 50% 来表示“可能会”,果然,最常使用 50% 的预测者准确率低于平均水平。

Forecasters commonly use 50 percent to represent “maybe,” and sure enough the forecasters who used 50 percent the most frequently were less accurate than the average.

超级预测者提供的预测比其他预测者更精细。他们更可能用 70%、75%、80%,甚至 70%、71%、72%,而不是 60%、70%、80%。这种精确并非为了炫耀:粒度更细的预测比粒度较粗的预测更准确。

Superforecasters provide more finely detailed forecasts than the other forecasters. Instead of 60, 70, 80 percent they are more likely to use 70, 75, 80 percent, or even 70, 71, 72 percent. This precision was not for show: the more granular forecasts were more accurate than the less granular ones.

哲学家们区分了两种情况:一种是不知道结果但可能性可知的情况(比如掷骰子);另一种是不知道结果且替代方案不可知的情况。超级预测者认识到,答案的前景越模糊,他们就越应该停留在“可能会”的区域。

Philosophers suggest a distinction between cases where you don’t know the outcome but the possibilities are knowable—the roll of a die for instance—and cases where you don’t know and the alternatives are unknowable.14 Superforecasters recognize that the cloudier the outlook for the answer, the more it benefits them to stay in the zone of “maybe.”

我们将在讨论工具时再回到概率问题,但再次强调反馈的价值仍是值得的。改善校准(即主观概率与客观概率的对齐)的最可靠方法之一,就是及时准确的反馈。气象预报员优于金融预报员的部分原因,在于他们能迅速看到自己的预测是否准确。

We will come back to probability when we discuss tools, but it is again worth underscoring the value of feedback. One of the surest ways to improve calibration, the alignment of subjective and objective probabilities, is through timely and accurate feedback. Part of the reason that weather forecasters are better than financial forecasters is that they quickly see whether their predictions are accurate.15

我们生活在一个动态的世界里,因此能够改变我们对概率评估的新信息时刻涌入。超级预测者最初的预测平均比普通预测者更准确,但他们也更频繁地更新自己的观点。这种更新需要开放的心态,但同时也伴随着反应不足或反应过度的风险。

We live in a dynamic world, so new information that should change our assessments of probabilities arrives all the time. Superforecasters have more accurate initial predictions on average than the regular forecasters do, but they also update their views more often. Such updating requires an open mind, but also comes with the risk of under- or overreacting.

泰洛克和加德纳提出了预测者对信息反应不足的三个原因。首先,有时我们太忙,新信息只是被我们忽略了。其次,我们可能偏离了最初的问题,而纠缠于一个更简单或略有不同的问题。因此,新信息看起来可能与我们脑子里的问题无关,尽管它实际上与手头的问题相关。最后,也是可能性最大的原因,是信念坚持。这通常伴随着确认偏误——主动寻找支持我们观点的信息,并摒弃与之相反的信息。

Tetlock and Gardner offer three reasons that forecasters underreact to new information. To start, sometimes we are so busy that novel information merely slips our attention. We also may take our eye off of the original question and dwell on a simpler or slightly different one. So the new information may not appear relevant to the question in our minds even though it is relevant to the question at hand. Finally, and probably most likely, is belief perseverance. This is typically accompanied by confirmation bias—actively seeking information that supports our view and dismissing views counter to it.16

但信息泛滥也可能让人反应过度。原因之一是,人们会把无关信息也纳入考量。最初判断或许建立在坚实推理之上,但随后却对与问题毫无关联的额外信息赋予权重。反应过度的第二个原因是缺乏定力。假设你买入一只共同基金,此前做过研究,确信基金经理能力出众、其策略长期会表现良好。若是业绩刚出现一点波动就卖出,那说明你缺乏定力。

But it’s also possible to overreact to new information. One reason is that we take into account irrelevant information. People may base their initial estimate on solid reasoning, but subsequently place weight on additional information that has no bearing on the issue at hand. A second reason for overreaction is a lack of commitment. Say you buy a mutual fund after having done research convincing yourself that the portfolio manager is skillful and that her strategy will do well over time. If you sell at the first bump in performance, you are showing a lack of commitment.

使用贝叶斯定理是更新概率的一种正式方法,但事实证明,超级预测者通常并不使用它,即使对那些精通该定理数学原理的人来说也是如此。没有简单的方法能正确更新观点,但我们知道超级预测者花了很多时间去思考如何把这个过程做好。

Use of Bayes’s Theorem is a formal way to update probabilities, but it turns out that the superforecasters generally don’t use it. This is true even for those steeped in the math of the theorem. There is no easy way to correctly update views, but we know that superforecasters spend a lot of time thinking about how to do it well.

我们认为大部分时间都在凭直觉行事。直觉系统倾向于依赖经验法则,也就是启发式方法。例如,可得性启发式表明,那些因生动性或近期性而更容易回想起来的事件,会被我们视为更频繁发生,而同样频率但不易记起的事件,则会被低估。

We think intuitively most of the time. Our intuitive system tends to use rules of thumb, or heuristics. For example, the availability heuristic suggests that we deem events that are easy to remember, because of vividness or recency, to be more numerous than events of equal frequency that are not as simple to recall.

其他启发式包括代表性和锚定效应。¹⁷

Other heuristics include representativeness and anchoring.17

启发式方法妙不可言,因为它能为我们节省大量时间。但同时也伴随着偏见,这些偏见可能损害我们的判断质量。例如,人们在听说一起航空事故后,会对飞行产生不合理的过度恐惧。

Heuristics are wonderful because they save us a great deal of time. But they also come with biases that can undermine the quality of our judgments. For example, people are unjustifiably more fearful of flying after having heard of an aviation accident.

超级预测者对这类偏见的觉察力高于常人,并有意识地加以管理。他们的思维方式和预测方法都有助于消除偏见。人与计算机的结合,则是另一条提升决策水平的路径。如果运用得当,在特定领域中,“人加机器”的输出可以超越单靠人类或单靠机器。

Superforecasters have an above-average awareness of these biases and try to manage them. Both the thinking styles and forecasting methods of superforecasters help address bias. The melding of humans and computers is another path to improved decisions. Done correctly, the output of man plus machine can exceed man or machine in certain domains.

戴维·费鲁奇是一位人工智能专家,曾负责打造沃森——这台计算机在电视答题节目《危险边缘》中击败了冠军肯·詹宁斯和布拉德·鲁特尔。他认为,即便在机器人崛起的时代,人的判断仍有一席之地。但他同时相信,计算机可以帮助克服人类的偏见。“所以我想要的是,让人类专家与计算机搭档,”费鲁奇说,“来克服人类认知的局限和偏见。”

David Ferrucci is an artificial intelligence expert who was in charge of creating Watson, the computer that beat champions Ken Jennings and Brad Rutter in the game show Jeopardy! He sees a role for human judgment even with the rise of the robots. But he also believes that computers can help overcome human bias. “So what I want is that human expert paired with a computer,” said Ferrucci, “to overcome the human cognitive limitations and biases.”

工作伦理。斯坦福大学心理学教授卡罗尔·德韦克(Carol Dweck)以其关于思维模式(mindset)的研究而闻名。她提出,根据人们对能力来源的潜在信念,可以将个体置于一个从“固定型思维”到“成长型思维”的连续谱系中。持固定型思维的人认为能力是天生的,因此无法改变。那些说“我就是数学不行”的人就属于固定型思维。而持成长型思维的人则认为能力是努力和付出的结果,因此可以随着时间的推移不断提升。

Work Ethic. Carol Dweck, a professor of psychology at Stanford University, is best known for her work on mindset, or a way of thinking. She suggests individuals can be placed on a continuum, with a “fixed mindset” at one extreme and a “growth mindset” at the other, based on their implicit belief about the source of ability. People with a fixed mindset believe that ability is innate and therefore can’t be budged. People who say, “I’m just bad at math,” have a fixed mindset. Those with a growth mindset believe that ability is the result of hard work and effort and therefore can improve over time.

超级预测者落在成长型思维这一端。他们认为总有改进余地,并不断寻找提升路径。在讨论自己某项实验结果时,德韦克指出:“只有具备成长型思维的人才会密切关注那些能拓展自己知识的信息。只有对他们而言,学习才是一项优先事项。”¹⁸

Superforecasters fall on the growth mindset side of the continuum. They believe that there is always room for improvement and seek ways to get better. In discussing the results of one of her experiments, Dweck noted, “Only people with a growth mindset paid close attention to information that could stretch their knowledge. Only for them was learning a priority.”18

超级预测者还有另一个品质——坚毅(grit),这个词由泰特洛克在宾夕法尼亚大学的另一位同事安吉拉·达克沃什推广开来。¹⁹ 坚毅是为长期目标服务的坚持不懈,意味着在实现目标的过程中具备克服失败和障碍的能力。

Another quality that the superforecasters have is grit, a term popularized by Angela Duckworth, another one of Tetlock’s colleagues at the University of Pennsylvania.19 Grit is perseverance in the service of long-term goals. It entails the ability to overcome failure and obstacles along the way of achieving an objective.

把成长型思维和坚毅结合起来,你就得到了个人发展与进步的绝佳配方。泰特洛克和加德纳把这种组合称为“永久测试版”。处于测试版的产品近乎完成,但仍留有改进空间。永久测试版意味着一种持续改进的渴望。图表 1 总结了一位超级预测者的综合画像。

Combine a growth mindset and grit and you have an outstanding formula for personal development and improvement. Tetlock and Gardner call the combination “perpetual beta.” A product in beta is nearly complete but has room for improvement. Perpetual beta suggests a desire for ongoing improvement. Exhibit 1 summarizes the composite portrait of a superforecaster.

表 1:典型超级预测者的综合画像

哲学取向

− 谨慎:没有什么是确定的

− 谦逊:现实是无限复杂的

− 非决定论:发生的事情并非命中注定,也不必然发生

能力与思维方式

− 积极主动的开放心态:信念是待检验的假设,而非待保护的珍宝

− 聪慧且知识渊博,具有“认知需求”:求知欲强,喜欢谜题和智力挑战

− 反思:内省且自我批判

− 擅数:与数字相处自如

预测方法

− 务实:不执着于任何想法或议程

− 分析:能够跳出眼前视角,考虑其他观点

− 蜻蜓之眼:重视多元观点并整合为自己的判断

− 概率思维:用多种可能性程度来评判

− 及时更新者:事实改变时,他们改变想法

− 优秀的直觉心理师:意识到需要检查思维中是否存在认知和情绪偏差

职业道德

− 成长型思维:相信自己能够进步

− 坚韧:无论花多长时间,都决心坚持下去

来源:菲利普·E·泰特洛克与丹·加德纳,《超级预测:预测的科学与艺术》(纽约:皇冠出版社,2015 年),第 191-192 页。经许可使用。

Exhibit 1: Composite Portrait of the Modal Superforecaster Philosophic Outlook − Cautious: Nothing is certain − Humble: Reality is infinitely complex − Nondeterministic: What happens is not meant to be and does not have to happen Abilities and Thinking Styles − Actively open-minded: Beliefs are hypotheses to be tested, not treasures to be protected − Intelligent and knowledgeable, with a “need for cognition”: Intellectually curious, enjoy puzzles and mental challenges − Reflective: Introspective and self-critical − Numerate: Comfortable with numbers Methods of Forecasting − Pragmatic: Not wedded to any idea or agenda − Analytical: Capable of stepping back from the tip-of-your-nose perspective and considering other views − Dragonfly-eyed: Value diverse views and synthesize them into your own − Probabilistic: Judge using many grades of maybe − Thoughtful updaters: When facts change, they change their minds − Good intuitive psychologists: Aware of the value of checking thinking for cognitive and emotional biases Work Ethic − A growth mindset: Believe it’s possible to get better − Grit: Determined to keep at it however long it takes Source: Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (New York: Crown Publishers, 2015), 191-192. Used by permission.

超级预测者的这些特质可能为招聘流程和绩效评估增添结构化价值,具有借鉴意义。此外,预测行业的组织领导者应当审慎思考如何营造有利于良好判断的环境。

These elements of a superforecaster may be valuable for adding structure to hiring processes and performance evaluation. Further, leaders of organizations in the forecasting business should give careful consideration to creating an environment conducive to good judgment.

Manage Interaction

Manage Interaction

IARPA 竞赛的目标是做出准确的预测,而 GJP 已经证明了自己能做到这一点。下一个问题是,团队合作能否进一步提高准确率。为了找到答案,第一年 GJP 的研究人员随机将部分预测者分配到团队中,并向他们提供了如何有效协作的建议。其他人则单独工作。这发生在科学家们识别出超级预测者之前。

The objective of the IARPA tournament was to make accurate forecasts, and the GJP had already proven that it could do that. The next question was whether working in teams would improve accuracy. To find out, in year one the GJP researchers randomly assigned some forecasters to work in teams and provided them with tips on how to work together effectively. Others were to work alone. This was before the scientists had identified the superforecasters.

结果很明确:团队平均比个人准确 23%。图 2 展示了前两年的统计结果。请记住,布莱尔评分越低越好。

The results were clear: teams were on average 23 percent more accurate than individuals. Exhibit 2 shows the statistical results for the first two years. Recall that a lower Brier score is better than a higher one.

图表 2:预测团队跑赢个人 0.15

Exhibit 2: Forecasting Teams Outperformed Individuals 0.15

平均标准化布里尔分数
   0.10
   0.05
   0.00
   -0.05
   -0.10
   -0.15
Mean Standardized Brier Score
   0.10
   0.05
   0.00
   -0.05
   -0.10
   -0.15

个人 团队 个人 团队 第一年 第二年 来源:基于 Barbara Mellers、Lyle Ungar、Jonathan Baron、Jaime Ramos、Burcu Gurcay、Katrina Fincher、Sydney E. Scott、Don Moore、Pavel Atanasov、Samuel A. Swift、Terry Murray、Eric Stone 和 Philip E. Tetlock 的研究,“Psychological Strategies for Winning a Geopolitical Forecasting Tournament”,《心理科学》,第 25 卷,第 5 期,2014 年 5 月,第 1106-1115 页。

Individual Team Individual Team Year 1 Year 2 Source: Based on Barbara Mellers, Lyle Ungar, Jonathan Baron, Jaime Ramos, Burcu Gurcay, Katrina Fincher, Sydney E. Scott, Don Moore, Pavel Atanasov, Samuel A. Swift, Terry Murray, Eric Stone, and Philip E. Tetlock, “Psychological Strategies for Winning a Geopolitical Forecasting Tournament,” Psychological Science, Vol. 25, No. 5, May 2014, 1106-1115.

注:误差线表示正负两个标准误差。

Note: Error bars represent plus and minus two standard errors.

大多数组织都依靠团队来完成工作。在包括投资管理在内的一些行业,趋势显示团队的使用正在增加。20 但研究也揭示了团队工作的利弊。有利之处在于,个体可以共享信息,而汇总往往能带来更准确的预测。不利之处在于,团队成员可能会产生偷懒的倾向,团队也可能陷入群体思维,从而无法捕捉认知多样性的价值。

Most organizations use teams to get their work done. In some sectors, including investment management, the trend shows an increasing use of teams.20 But the research also reveals that there are pros and cons to working in teams. The pros are that individuals can share information, and aggregation tends to lead to more accurate forecasts. The cons are that team members may be tempted to loaf, and the team may fall into groupthink, failing to capture the value of cognitive diversity.

第二年,研究人员组建了“超级预测者”团队。每个团队有十几名成员,但大部分工作由五到六个人完成。团队同样接受了协作培训,由于成员们无法面对面交流,项目协调员为他们创建了沟通论坛。

In year two, the researchers created teams of superforecasters. Each team had a dozen members, but about five or six individuals did most of the work. Once again, the teams were trained on how to work together, and because the members did not meet face-to-face the project coordinators created forums for them to communicate.

结果令人惊叹。那些在第一年就成功跻身超级预测者行列的人,在第二年作为团队一员时,平均准确率提升了 50%。第三年情况类似。(见图表 3。)事实上,这些超级团队的表现比预测市场还要好 15% 到 30%,而超越预测市场本身就是一个极高的标准。

The outcome was remarkable. Those who were good enough to achieve superforecaster status in year 1 were 50 percent more accurate, on average, in year 2 as part of a team. Year three was more of the same. (See Exhibit 3.) Indeed, these superteams were 15 to 30 percent better than prediction markets, a high standard to exceed.

图表 3:超级预测者在团队中表现更佳——0.1 个标准差

Exhibit 3: Superforecasters Are Even Better in Teams 0.1

Mean Standardized Brier Score
0.0
-0.1
-0.2
-0.3
-0.4
第 1 年第 2 年第 3 年
顶级团队
Mean Standardized Brier Score
   0.0
   -0.1
   -0.2
   -0.3
   -0.4
   Year 1   Year 2   Year 3
   Top-Team

超越所有其他人个人来源:基于芭芭拉·梅勒斯、埃里克·斯通、特里·默里、安吉拉·明斯特、尼克·罗尔堡、迈克尔·毕晓普、陈伊娃、约书亚·贝克、侯元、迈克尔·霍洛维茨、莱尔·昂加尔和菲利普·泰特洛克合著论文《识别与培养超级预测者作为改进概率预测的方法》,载于《心理科学视角》第 10 卷第 3 期,2015 年 5 月,第 267–281 页。

Supers All Others Individuals Source: Based on Barbara Mellers, Eric Stone, Terry Murray, Angela Minster, Nick Rohrbaugh, Michael Bishop, Eva Chen, Joshua Baker, Yuan Hou, Michael Horowitz, Lyle Ungar, and Philip Tetlock, “Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions,” Perspectives on Psychological Science, Vol. 10, No. 3, May 2015, 267–281.

注意:误差棒代表正负一个标准误。

Note: Error bars represent plus and minus one standard error.

这并不是说一切进展顺利。起初,这些超级预测者很谨慎,对于那些可能被认为过于挑剔他人判断的评论,他们会加以克制。而且,并非所有超级预测者都拥有出色的社交技巧。不过,在大多数情况下,团队成员们找到了相互进行建设性交锋的方法,这使得团队能够提升表现。在第二年和第三年之后,超级预测者们被邀请进行面对面聚会,其中许多人报告说,增加的人际维度与责任感对工作很有帮助。

This is not to say that all went smoothly. Initially, the superforecasters were cautious, restraining comments that would be perceived to be too critical of the assessments of others. And not all of the superforecasters had great social skills. In most cases, however, the teammates figured out ways to engage one another in constructive confrontation that allowed the groups to improve performance. The superforecasters were invited to get together in person after years two and three, and many of them reported that the added human dimension and sense of commitment were helpful.

尽管超级团队取得了成功,但 GJP 的研究人员对这种模式在企业环境中的适用性持保留态度。将个人认定为“超级”、让公司不同部门的员工协同工作以及人际互动中的动态变化,都可能产生摩擦。

Notwithstanding the success of the superteams, the GJP researchers have reservations about how well the formula would apply in a corporate setting. Identifying individuals as “super,” getting employees from different parts of the firm to work together, and interpersonal dynamics might all create friction.

主要教训是,在管理得当的情况下,一个多样化群体——尤其是那些具备超级预测者特征的群体——内部的互动可以非常有效。追求多样性的公司,也应准备好对员工进行多样性管理方面的培训。

The main lesson is that interaction among a diverse group, especially those with the profile of a superforecaster, can be very effective if managed properly. Companies that seek diversity should also be prepared to train employees in how to manage diversity.21

Train Effectively

Train Effectively

GJP 的一个显著发现是:相对较少的训练——大约一小时——就能将判断结果提升 10%。(见图表 4。)对于个人来说,训练的重点在于提升概率推理能力并消除认知偏差。一个例子就是我们倾向于使用“内部视角”还是“外部视角”。采用内部视角时,我们倾向于依赖自身掌握的信息和感知。外部视角则将问题视为一个更大参照类别中的实例,并依赖过去发生的基准概率。虽然良好的判断中两种视角都很重要,但心理学家已证明,我们通常过度依赖内部视角。

One remarkable finding from the GJP is that relatively little training, about an hour, can improve results by 10 percent. (See Exhibit 4.) For individuals, the training focused on sharpening probabilistic reasoning and removing cognitive biases. One example is how we tend to use the “inside” versus the “outside” view. With the inside view, we tend to rely on our own information and perception. The outside view considers a problem as an instance of a larger reference class and appeals to the base rate of past occurrence. While both are important in good judgment, psychologists have shown that we commonly rely too much on the inside view.

让人们意识到外部视角有助于减少这种偏见。

Making individuals aware of the outside view can help reduce this bias.

展示 4:预测者从去偏训练中获益 0.10

Exhibit 4: Forecasters Benefit from Debiasing Training 0.10

标准化平均布里尔评分
0.05
0.00
-0.05
-0.10
-0.15
无训练无训练
第一年第二年
Mean Standardized Brier Score
   0.05
   0.00
   -0.05
   -0.10
   -0.15
   None Training   None Training
   Year 1   Year 2

资料来源:芭芭拉·梅勒斯、莱尔·昂加尔、乔纳森·巴伦、海梅·拉莫斯、布尔楚·古尔凯、卡特里娜·芬彻、悉尼·E·斯科特、唐·摩尔、帕维尔·阿塔纳索夫、塞缪尔·A·斯威夫特、特里·默里、埃里克·斯通和菲利普·E·泰特洛克,“地缘政治预测锦标赛中的心理制胜策略”,《心理科学》,第 25 卷,第 5 期,2014 年 5 月,第 1106-1115 页。

Source: Based on Barbara Mellers, Lyle Ungar, Jonathan Baron, Jaime Ramos, Burcu Gurcay, Katrina Fincher, Sydney E. Scott, Don Moore, Pavel Atanasov, Samuel A. Swift, Terry Murray, Eric Stone, and Philip E. Tetlock, “Psychological Strategies for Winning a Geopolitical Forecasting Tournament,” Psychological Science, Vol. 25, No. 5, May 2014, 1106-1115.

注:误差条表示正负两个标准误。

Note: Error bars represent plus and minus two standard errors.

针对团队而言,培训的内容是如何协作。目标是在冲突与和谐之间找到平衡。冲突太多,团队动力就会瓦解,没有人愿意互动。和谐太多则会导致虚假共识,甚至群体迷思。恰当的平衡,是营造一种建设性批评的氛围。

For groups, the training addressed how to work together. The goal is to strike a balance between conflict and harmony. Too much conflict and the group dynamics break down. No one wants to interact. Too much harmony leads to a false consensus, or even groupthink. The right balance is an atmosphere of constructive criticism.

对于精英预测者的过度权重分配,或是将估计值推向极端化。

Overweight Elite Forecasters or Extremize Estimates

获取群体智慧最简便的办法,就是把许多持有不同观点的人给出的估算汇总起来。例如,你可以让一大群人观察一个装满软糖豆的罐子,请每个人估算总数。群体预测的结果会优于个人预测的平均值,而且通常非常接近实际的软糖豆数量。

The simplest way to capture the wisdom of crowds is to aggregate the estimates of a large number of people with diverse views. For example, you can let lots of people examine a jar filled with jelly beans and ask each of them to provide an estimate of the total. The collective prediction will beat the average individual predictions and will generally be very close to the actual number of jelly beans.

在 GJP 项目中工作的统计学家找到了两种提高预测质量的方法。22 第一种方法是给予超级预测家更大的权重。这背后的直觉很直接:给那些预测更准确的人更大的发言权。

The statisticians working on the GJP figured out two ways to improve the quality of forecasts. 22 The first is to place more weight on what the superforecasters say. The intuition behind this is straightforward; give those who predict more accurately a greater voice.

第二种是应用一个算法来“极端化”答案。想象一下,有五个人持不同观点,但都估计某个结果的概率为 75%。极端化算法会将这个概率推向接近 100%。同样,一个给出低概率的群体估计也会被调整到接近 0%。算法将群体预测推向 0 或 100 的程度,取决于预测者群体的多样性和成熟度。

The second is the application of an algorithm to “extremize” answers. Imagine five individuals with different points of view who all estimate the probability of an outcome to be 75 percent. The extremizing algorithm would push that probability closer to 100 percent. Likewise, a group estimate with a low probability would be adjusted closer to zero percent. How far the algorithm pushes the aggregate forecast toward 0 or 100 is a function of the diversity and sophistication of the pool of forecasters.

这个思路是要捕捉未共享的信息。以下是一种理解方式:我们那些给出 75% 概率的多元预测者,每个人既使用他们共有的信息,也使用他们独有的信息。如果每个人都知道全部信息,会发生什么?那会增强他们的信心,使他们的集体估计趋近于 100%。因此,极端化算法所捕捉的,正是如果多元个体能将自己所有信息与群体共享时会发生的情况。

The idea is to capture unshared information.23 Here’s one way to think about it. Each of our diverse forecasters who come up with a probability of 75 percent uses some information that’s common to all of them and some that’s unique to them. What would happen if each of them knew all of the information? It would strengthen their confidence, moving their collective estimate toward 100 percent. So the extremizing algorithm captures what would happen if diverse individuals could share all of their information with the group.

A Superforecaster’s Tools

A Superforecaster’s Tools

我们已经讲了超级预测者需要具备什么条件,以及准确预测包含哪些要素。但略过了一些做出好预测所必需的核心工具。下面就来介绍这些工具,包括如何界定一个好问题、如何记分,以及预测的几种方法。

We have covered what makes for a superforecaster and the elements of an accurate prediction. But we have glossed over some of the essential tools necessary to come up with good forecasts. We now cover those tools, including defining a good question, keeping score, and approaches to forecasts.

好问题。要得出有用的答案,第一步就是提出好问题。一个好问题,应该具有在特定时间框架内明确的预期结果,并且要涉及与这个世界切实相关的议题。我们来逐一审视这些要素。

Good Questions. The first step in coming to useful answers is to ask good questions. A good question has a clear outcome within a specified period of time and addresses an issue that is relevant to the world. Let’s look at each of these components.

因为记分对培养预见能力至关重要,所以问题必须有清晰的答案。

Because keeping score is essential to developing the skill of foresight, questions must have clear answers.

以天气这个简单的例子来思考。天气预报员预测的是气温、降水量和风力。

Consider the simple example of weather. Weather forecasters predict temperature, precipitation, and winds.

他们可以对每一项进行衡量,并将预测结果与实际结果进行对比。同样,GJP 和情报机构提出的问题也应该有明确的答案。

They can measure each of these and compare the prediction to the outcome. Likewise, questions in the GJP and those from the intelligence community should have clear answers.

问题还必须设定明确的时间框架。泰特洛克此前关于专家判断的研究结果并不理想,而最近预测竞赛的结果则描绘了一幅更为乐观的图景——两者之间的一个关键区别就在于时间框架。此前的研究中,问题跨度是 3 到 5 年,而在 IARPA 竞赛中,问题平均跨度只有几个月到一年。要在 3 到 5 年的时间跨度里预测政治、经济和社会领域的结果,对任何人来说都极其困难。时间框架更短的预测则更具可行性,同时对商业和政策决策仍有参考价值。

Questions must also have a set time frame. One key difference between Tetlock’s prior research on expert judgment, which showed poor results, and the latest results from the forecasting tournament, which paint a more optimistic picture, is the time frame. In the prior research, the questions went out three to five years whereas in the IARPA tournament they average about a few months to a year. Predicting outcomes in the political, economic, and social realms over three to five years is very hard for anyone. Predictions in shorter time frames are more feasible and yet still relevant for business and policy decisions.

开放式预测的价值有限。泰特洛克和加德纳讨论了一封 2010 年 11 月写给时任美联储主席本·伯南克的信。这封信由近二十位经济学家、投资者和评论员联署,建议量化宽松应当“重新考虑并停止”,并指出存在“货币贬值和通货膨胀”的风险。2014 年秋季,彭博社的记者联系了部分联署人,发现“所有作出回应的人都坚持信中的内容”。由于没有设定时间框架或客观的评分方法,这封信无法让人对预测进行计分。

Open-ended predictions are of limited value. Tetlock and Gardner discuss a letter sent in November 2010 to Ben Bernanke, then chairman of the Federal Reserve. Signed by nearly two dozen economists, investors, and commentators, the letter suggested that quantitative easing should be “reconsidered and discontinued” and noted the risk of “currency debasement and inflation.” In the fall of 2014, reporters at Bloomberg contacted some of the signees and found that “all of those who commented stood by the letter’s contents.” With no time frame or objective way to score the prediction, the letter provides no way to keep score.

最后,好的问题还要贴合实际。有些问题可能是一组问题的子集,借助它们你可以评估一个更宏大但更棘手的议题。泰特洛克和加德纳提出,一个有用的问题应该能通过“拍脑门测试”:“当时间过去后你重读这个问题,会拍着脑门说,‘我当初怎么就没想到这一点!’”。

Finally, good questions are also relevant. Some questions may be a subset of a group of questions that allow you to assess a larger, but more difficult, issue. Tetlock and Gardner suggest a useful question should pass the smack -the-forehead test: “when you read the question after time has passed, you smack your forehead and say, ‘If only I had thought of that before!’”

以下是预测锦标赛中的一个问题示例。请注意,它具备明确的结果、指定的时间范围和相关性:24

Here’s an example of a question from the forecasting tournament. Note that it has a clear outcome, specified time period, and relevance:24

“意大利的西尔维奥·贝卢斯科尼会在 2012 年 1 月 1 日之前辞职、连任失败或失去信任投票,或以其他方式离任吗?”

“Will Italy’s Silvio Berlusconi resign, lose reelection/confidence vote, or otherwise vacate office before 1 January 2012?”

泰特洛克和加德纳还提出了另一个关于问题的观点。他们认为,提出好问题的技能可能不同于回答问题的技能。超级提问者未必是超级预测者。有远见的组织应当努力培养这两种能力。

Tetlock and Gardner make an additional point on questions. They suggest that the skills in asking good questions may be different than the skills in answering them. The superquestioners might not be the same as the superforecasters. Thoughtful organizations should seek to cultivate skills in both capacities.

计分。我们已经提到过布莱尔分数(Brier score),这是心理学家用来衡量结果的主要方法。如果你有一个带有具体概率的预测和一个结果,你就能对这些预测进行评分。由于布莱尔分数反映了预测与结果之间的距离,数字越低越好。超级预测者就是在这场竞赛中拿到最低布莱尔分数的人。

Keeping Score. We have already mentioned the Brier score, the main method that psychologists use to measure results. If you have a prediction with a specific probability and an outcome, you are in a position to grade the predictions. Since Brier scores reflect the distance from a forecast to the outcome, lower figures are better. Superforecasters were those in the tournament who earned the lowest Brier scores.

保持记录至关重要,因为这会提供反馈,从而带来学习机会。预测者希望在两个方面有所提升。一个方面叫作“校准”,意思是你的预测与结果保持一致。例如,如果你说某些事件发生的概率为 40%,而它们实际上在 40% 的情况下发生了,那么你的校准就做得很好。

Keeping score is crucial because it provides feedback and therefore an opportunity to learn. Forecasters want to improve in two ways. One way is called “calibration,” which means that your forecasts line up with the outcomes. For example, you are well calibrated if you say that certain events will occur with a 40 percent probability and they actually happen 40 percent of the time.

另一个提升的方法是作者所谓的“判定力”。判定力指的是,当你确信某件事不会发生时,它确实没有发生;或者当你确信它会发生时,它真的发生了。这是信念程度的度量。良好的校准能力与判定力相互关联,但二者截然不同。你需要在这两方面都精进技能。

Another way to improve is what the authors call “resolution.” Resolution means that when you are sure something is not going to happen, it doesn’t happen, or when you’re sure it will happen, it does. It’s a measure of conviction. Good calibration and resolution are correlated, but they are distinct. You want to sharpen your skills in both ways.

泰特洛克很喜欢一个说法:“模糊的措辞减缓了学习周期。” 在日常对话中,我们会使用大量模棱两可的词语——也许、可能、大概、有可能会发生。有一个著名的例子:1951 年,美国情报界发布了一份报告,称苏联入侵南斯拉夫是“严重的可能性”。当时在中央情报局供职的耶鲁历史学教授谢尔曼·肯特,询问他团队中的几位成员,“严重的可能性”究竟是什么意思。尽管他们都同意在报告中使用这个措辞,但一位分析员认为这意味着 80% 的概率,另一位则认为只有 20%。

Tetlock is fond of the phrase “vague verbiage slows learning cycles.”26 In our day-to-day conversations we use lots of phrases—maybe, possible, probable, might happen—that are ambiguous. In one famous illustration, in 1951 the U.S. intelligence community produced a report saying that the Soviet Union’s invasion of Yugoslavia was a “serious possibility.” Sherman Kent, a professor of history at Yale who was then serving at the Central Intelligence Agency, asked some members of his team what “serious possibility” meant. Even though they all agreed to use the term in their report, one analyst said it meant an 80 percent probability and another said 20 percent.

这则故事的教训是:只要有可能,就要使用具体的概率和时间跨度,并且对这些预测进行追踪记录。数值概率消除了肯特所面临的被误解或曲解的风险,并为预测者提供了改善所需的反馈。

The lesson from this story is to use specific probabilities and time horizons whenever possible and to keep track of those forecasts. Numerical probabilities dismiss the risk of misinterpretation or misunderstanding that Kent faced and provide forecasters with the feedback they need to improve.

在预测方法的研究中,泰特洛克及其同事发现,超级预测者在处理任务时采用的一些方法,或许在预测竞赛的整洁范围之外也有用武之地。

Approaches to Forecasts. In studying how superforecasters approach their task, Tetlock and his colleagues noticed a few methods that may be useful beyond the tidy confines of the forecasting tournament.

第一个理念是问题分类(question triage)。在战争等医疗紧急情况下,医生和护士会根据伤情对伤员进行分类。这个过程叫做分类。需要立即救治的患者优先级高于预期能存活的患者和没有生还希望的患者。同样,你也可以对问题进行分类。有些问题太简单,有些则太难。优先级应该给那些介于这两个极端之间的、属于“付出努力回报最大”的问题。

The first is the idea of question triage. In times of medical emergency such as war, doctors and nurses sort casualties based on their injuries. This process is called triage. Patients who require immediate care receive priority over those who are expected to live and those who have no chance of survival. Similarly, you can sort questions. Some are too easy and others too hard. The priority should go to questions that are between those extremes where “effort pays off the most.”

超级预测者还会运用特定技巧来回答一些更大、但可回答的问题。其中一种技巧源自芝加哥大学教授恩里科·费米(Enrico Fermi),他曾参与曼哈顿计划并荣获诺贝尔物理学奖。这本质上是一种信封背面的速算方法。费米用这种方法估算原子弹试验的威力,但你可以将这一技巧用于任何有答案的问题。

Superforecasters also use specific techniques to answer some bigger, but answerable, questions. One is associated with Enrico Fermi, a professor at the University of Chicago who worked on the Manhattan Project and won the Nobel Prize in physics. It is effectively a back-of-the-envelope calculation. Fermi used the approach to estimate the strength of an atomic bomb test, but you can use the technique for any question for which there is an answer.

在《超预测》一书中,泰洛克和加德纳用一个经典的费米问题举例:芝加哥有多少名钢琴调音师?方法是将这个大问题拆解成一系列更小、更容易回答的问题。你可以先估算芝加哥的人口,评估家庭数量,判断每户家庭拥有钢琴的数量,再将学校和宗教场所考虑进去,确定钢琴需要调音的频率,并思考调音一台钢琴需要多久,一名调音师一年能工作多少小时。最终得出,芝加哥大约有 250 到 300 名钢琴调音师。

In Superforecasting, Tetlock and Gardner use a popular example of a Fermi problem: How many piano tuners are there in Chicago? The tactic is to break the big question into a series of smaller, and easier, questions. You might estimate the population of Chicago, assess the number of households, judge how many pianos there are per household, consider the number of schools and places of worship, determine how often pianos need to be tuned, and contemplate how long it takes to tune a piano and how many hours a piano tuner might work. It turns out there are about 250-300 piano tuners in Chicago.27

另一种技术叫作“问题聚类”。其核心理念在于:我们想回答的大问题——“朝鲜会发动战争吗?”——过于宏大,而小问题——“朝鲜会在某个日期前发射多级火箭吗?”——又过于微小。泰特洛克和加德纳建议,将一系列小问题聚合成群,每个小问题你都能持续更新判断,这些判断组合起来就能为回答那个更大的问题提供洞见。他们用的比喻是“

Another technique is called question clustering. The idea is that the big question we want to answer—“Will North Korea initiate a war?”—is too big, and the small questions—“Will North Korea launch a multistage rocket by some date?”—are too small. Tetlock and Gardner suggest that a cluster of small questions, each of which you can update, provide insight into answering the bigger question. They use the metaphor of the

点彩画法的绘画技巧是在画布上添加一个个色点。单独一个点本身没什么意义,但汇聚在一起就能描绘出一幅生动感人的画面。

painting technique, pointillism, which consists of adding dots to the canvas. No dot by itself means much, but together they paint an evocative picture.

为了评估朝鲜发动侵略的可能性,他们建议你可以问一些关于导弹发射、核试验、网络攻击和火炮射击的问题。这些概率模式为处理一些重大问题提供了一条路径。同样值得注意的是,提出问题的人和回答问题的人不一定是同一个人。

To gauge the likelihood of North Korean aggression, they propose that you might ask questions about missile launches, nuclear tests, cyber-attacks, and artillery shelling. The patterns of probabilities provide a path to tackling some of the big questions. Here again, it’s good to remember that those who come up with the questions and those who answer need not be the same.

Leadership

Leadership

泰特洛克和加德纳指出,如果你让人们列举强大领导者的特质,你会听到诸如自信、果断、有远见这样的形容词。但正如我们所见,这些描述并不太符合超级预测者的特点。领导者应该始终处于行动模式,而超级预测者则似乎始终处于学习模式。

Tetlock and Gardner suggest that if you ask people to list the qualities of a strong leader, you will hear adjectives such as confident, decisive, and visionary. But as we have seen, those descriptions don’t fit the superforecasters very well. Leaders are supposed to be in perpetual action mode, while superforecasters seem to be in perpetual learning mode.

作者引用了普鲁士军队参谋长赫尔穆特·冯·毛奇的智慧,他执掌普鲁士军队长达三十年,被许多人视为现代战场指挥方法的开创者。“战争之中,一切皆不确定,”毛奇说,他认识到知识的局限。同时,他也承认灵活性的必要性,指出:“任何作战计划,在与敌军主力首次交锋之后,都无法再笃定地继续执行。”

The authors refer to the wisdom of Helmuth von Moltke, the chief of staff of the Prussian Army for three decades and considered by many to be the inventor of the modern method of directing troops in the field. “In war, everything is uncertain,” said Moltke, recognizing the limits of knowledge. And acknowledging the necessity of flexibility, he observed, “No plan of operations extends with certainty beyond the first encounter with the enemy’s main strength.”

毛奇的观点——最终被德军采纳——在于领导者必须思考,并在必要时调整行动方向。命令可以被质疑,甚至批评,只要有更好的方式。其要义是主动保持思想开放。一本德国军事手册指出:“领导的艺术,在于及时认清形势,以及何时需要做出新的决策。”

Moltke’s point, which was eventually embraced by the German army, is that leaders must think and change their course of action as necessary. Orders could be questioned, even criticized, if there was a better way. The message was to be actively open-minded. A German military manual notes, “The art of leadership consists of the timely recognition of circumstances and of the moment when a new decision is required.”

戴维·彼得雷乌斯,退役的美国陆军四星上将,在当代语境中拥抱了毛奇的许多主题。他支持派遣自己的军官到顶尖大学攻读研究生课程,以便听取与他们自身观点不同的见解。这创造了关键的思维灵活性,听起来很像超级预测者所做的事情。

David Petraeus, a retired four-star U.S. Army general, embraced many of Moltke’s themes in a contemporary context. He supported sending his officers to top universities to pursue graduate studies in order to hear points of view that were different from their own. This creates vital mental flexibility and sounds a lot like what superforecasters do.

但彼得雷乌斯也深知,领导意味着既要思考又要行动。他说,一个领导者“需要判断出正确的举措,然后果断执行”。谦逊应使领导者审慎反思自身的所作所为,而自信则赋予一个人付诸行动的力量。

But Petraeus also recognized that leading requires thinking and doing. He said that a leader “needs to figure out what’s the right move and then execute it boldly.” Humility should make a leader think carefully about what he or she is doing, and confidence should give an individual the strength to act.

归根结底,我们希望领导者自信、果断且富有远见。但《超级预测》给我们的启示是:思考是行动的重要前提。此外,情况会发生变化,及时更新判断、适当调整战略方向至关重要。

Ultimately, we want our leaders to be confident, decisive, and visionary. But the lesson of Superforecasting is that thinking is an important precursor to action. Further, circumstances change and it is vital to update views and change strategic course appropriately.

Doubts

Doubts

菲尔·泰特洛克在其领域内备受尊敬,并且和他研究的超级预测者一样,始终在寻求进步。他的两位友人与同行——丹尼尔·卡尼曼与纳西姆·塔勒布——对泰特洛克所识别的预测能力提出了一些质疑。

Phil Tetlock is highly respected in his field and, like the superforecasters he studies, is in constant search of improvement. Two of his friends and colleagues, Daniel Kahneman and Nassim Taleb, offer some challenges to the skill Tetlock has identified in forecasting.

卡尼曼的挑战涉及“范围不敏感”(scope insensitivity)。这个概念是说,我们面对的问题常常会引发某种情绪感受。科学家可以用金钱来衡量这种情绪影响。

Kahneman’s challenge relates to “scope insensitivity.” The idea is that the questions we face often evoke a feeling of emotion. Scientists can measure that affect with money.

例如,研究人员可能会问你,愿意出多少钱来拯救 2000 只候鸟免于因石油泄漏、湿地破坏或除草剂和农药残留等危害而死亡。受访者给出的答案是大约 80 美元。你对这种情况感到难过,愿意支付一定金额来提供帮助。

For example, researchers might ask you how much money you would be willing to pay to save 2,000 migratory birds from dying from hazards, including oil spills, wetlands destruction, or residue from herbicides and pesticides. Subjects say about $80. You feel bad about the situation and are willing to pay some amount of money to help.

问题在这里:当研究人员询问被试,他们愿意花多少钱来拯救 2 万只鸟时,回答是 78 美元。而当问及拯救 20 万只鸟时,回答是 88.28 美元。这个金额反映的是人们对场景的感受,而不是他们愿意为拯救每只鸟支付多少钱。这里根本没有成本效益分析。被试用的是系统 1 思考,结果就是对规模无感。

Here’s the issue: When researchers asked subjects how much they would be willing to pay to save 20,000 birds, the answer was $78. And for 200,000 birds the response was $88.28 The dollar sum reflects how people feel about the scenario, not how much they would be willing to pay to save each bird. There is no cost-benefit analysis. Subjects think using System 1 and as a result are insensitive to the scope.

卡尼曼认为,一种类似的范围不敏感现象可能也会影响 GJP 的预测者。具体来说,他提出范围不敏感可能与时间框架有关。例如,当被问及某个独裁政权倒台的可能性有多大时,人们会聚焦于概率本身,却未能恰当地考虑时间范围。

Kahneman thought that a similar type of scope insensitivity might affect the GJP forecasters. Specifically, he thought scope insensitivity would have to do with the time frame. For example, if you are asked how likely it is that a particular dictatorial regime will fall you will focus on the probability and fail to properly consider the time frame.

经过计算,GJP 的科学家们发现,确实有许多预测者存在范围不敏感的问题。

Running the numbers, the GJP scientists found that indeed many of the forecasters were scope insensitive.

当被问及叙利亚总统巴沙尔·阿萨德政权倒台的可能性时,常规预测者给出的概率是:3 个月内发生的可能性为 40%,6 个月内发生的可能性为 41%。他们对两倍于第一个时间段给出的概率几乎相同。

When asked the probability that the regime of Syrian President Bashar al-Assad would fall, regular forecasters assigned a 40 percent probability that it would happen within 3 months and a 41 percent chance within 6 months. They assigned essentially the same probability to a period twice as long as the first.

但超级预测者的表现要好得多。他们给出的概率是:3 个月内为 15%,6 个月内为 24%。这谈不上完美的灵敏度,但比普通预测者要准得多。泰特洛克和加德纳将这一结果解读为:超级预测者比大多数预测者更善于调动系统 2 思维。事实上,他们认为超级预测者已将这种思维方式内化到了近乎自动化的程度。

But the superforecasters did much better. Their probabilities were 15 percent for 3 months and 24 percent for 6 months. That’s not perfect sensitivity, but it is a lot closer than the regular forecasters. Tetlock and Gardner interpret this result as the ability of superforecasters to engage System 2 thinking more readily than the majority of forecasters. Indeed, they argue that superforecasters have internalized this way of thinking to the point that it has become automatic.

纳西姆·塔勒布让“黑天鹅事件”这个概念流行起来——它来得突然,影响巨大,事后总能被解释得头头是道。泰特洛克和加德纳说,“塔勒布坚称,黑天鹅,也只有黑天鹅,决定着历史的走向。”如果真是这样,那么 IARPA 的那场预测竞赛就没什么价值了。

Nassim Taleb has popularized the concept of a black swan event, which comes as a surprise, is consequential, and is explained after the fact. Tetlock and Gardner say that “Taleb insists that black swans, and black swans alone, determine the course of history.” If so, IARPA’s forecasting tournament is of little value.

泰特洛克和加德纳提出了一些观点来反驳这种看法。第一个观点与“黑天鹅”的定义有关。如果我们知道结果分布的样子,即使包含极端事件,那么我们所面对的也是“灰天鹅”,而非黑天鹅。科学家在归类某些现象(包括地震、恐怖行为以及电网故障)的分布方面做了大量工作。虽然预测某个具体结果非常困难,但科学家对这些系统的行为模式有总体了解。塔勒布将这类情况称为“可建模的极端事件”,它们涵盖了我们关注的大部分内容。29

Tetlock and Gardner offer some thoughts to counter this view. The first has to do with the definition of a black swan. If we know what a distribution of outcomes looks like, even those with extreme events, then we are dealing with “gray swans,” not black swans. Scientists have done a lot of work classifying the distributions of certain phenomena, including earthquakes, terrorist acts, and power-grid failures. While predicting a specific outcome is very difficult, scientists have a general understanding of how these systems behave. Taleb calls these “modelable extreme events” and they capture much of what is of interest to us.29

黑天鹅和灰天鹅是极不可能发生的结果。因此,单凭预测竞赛的数据,根本不足以断定超级预测者是否擅长识别它们。在

Black swans and gray swans are highly improbable outcomes. As a result, there are simply not enough data from the forecasting tournament to conclude that superforecasters are either good or bad at spotting them. In

事实上,即便这场比赛持续数十年,也很难知道结果。因此,那些寻求对黑天鹅或灰天鹅事件进行准确预测的人,只能另寻他处了。

fact, even if the tournament ran for decades it would be hard to know. So those seeking accurate forecasts of black or gray swan events will have to look elsewhere.

但事实是,这场预测锦标赛确实涉及了许多与高管、投资者和决策者相关的问题。黑天鹅事件固然对世界影响巨大,但众多小事件的累积也同样举足轻重。思考未来并非要么是黑天鹅,要么是零;还有很多空间来考虑影响较小的结果。

But the fact is that the forecasting tournament does address a lot of issues that are relevant for executives, investors, and policymakers. While black swans have a large impact on the world, so do the accumulation of lots of small events. Thinking about the future is not black swan or nothing; there is plenty of room to consider outcomes that have smaller impact.

泰特洛克、卡尼曼和塔勒布在一件事上意见一致:对未来许多年做预测非常困难。超过一定时间跨度后,可预测性存在严重局限。根本问题在于,短于一年的精准预测是否有价值。大多数专业人士会斩钉截铁地回答:有。

One point on which Tetlock, Kahneman, and Taleb agree is that forecasting many years into the future is very difficult. There are severe limits on predictability beyond a certain period of time. The essential issue is whether there is value in sharp forecasts for horizons less than a year. Most professionals would answer with a resounding yes.

Summary

Summary

人类有一种根深蒂固的欲望,想要预知未来。对预测的需求,自然会催生一批预言家、权威人士和专家来供给“接下来会发生什么”的答案。但我们通常不去衡量这些预测的质量——而一旦有人真的去衡量,结果往往乏善可陈。

Humans have a deep-seated desire to anticipate the future. The demand for forecasts is met by a supply of seers, pundits, and experts expounding on what will happen next. But we generally don’t measure the quality of predictions, and when it has been done the results are unimpressive.

超级预测力证明,预测或许并非徒劳无益。美国情报界赞助的“精准预测项目”竞赛表明,有些预测者不仅表现出色,而且持续稳定。运用心理学最佳成果与精细测量法,该项目团队已能够为从事预测工作的人提供关键课程。以下为主要结论:

Superforecasting shows that prediction may not be so futile after all. The Good Judgment Project, part of a forecasting tournament sponsored by the U.S. intelligence community, revealed that some forecasters are not only good but consistently good. Using the best ideas from psychology and careful measurement, the GJP team has been able to provide essential lessons for anyone in the prediction business. Here are some of the main conclusions:

预测技能确实存在。研究人员发现,在预测群体中,有很小一部分人的准确率远高于平均水平,并且这种优势持续稳定。他们之所以能找出这些超级预测者,是因为他们接纳了大量预测者样本,提出了时间跨度不超过一年的问题,并持续跟踪每个人的回答记录。

Forecasting skill exists. The researchers found that a small percentage of the forecasting population were much more accurate than average and consistently so. They were able to find the superforecasters because they welcomed a large sample of forecasters, asked questions with time frames of a year or less, and kept track of the responses.

思维方式至关重要。对超级预测者的深入分析显示,他们确实聪明,但并非天赋异禀。让他们与普通预测者拉开差距的,是他们的思考方式。

Way of thinking is vital. Closer analysis of the superforecasters shows that they are bright, but not extraordinarily so. What distinguishes them from regular forecasters is the way they think.

超级预测者心态开明、认同非决定论、智力上谦逊、善于量化思考、会基于新信息认真调整判断,而且工作极为勤奋。

Superforecasters are actively open-minded, nondeterministic, intellectually humble, numerate, thoughtful updaters, and hard working.

团队。当超级预测者相互交流时,他们的预测会变得更准。但团队协作有利有弊:好处包括获取更多信息以及发挥整合聚合的力量;弊端则是社会懈怠和群体思维的风险。表现最佳的团队,其成员都接受过如何协作的指导和训练。

Teams. When superforecasters interact with one another, their predictions improve. But there are pros and cons to working in teams. The pros include more information and the ability to harness the power of aggregation. The cons are social loafing and the risk of groupthink. Teams that do the best have members who have been given instruction and training on how to work together.

训练。对去偏预测和有效协作的方法进行训练,能够改善结果。

Training. Training in methods to de-bias forecasts and to collaborate effectively improves outcomes.

恰当的培训应具备有价值的内容以及能够立即实施的过程。无法付诸实践的培训纯属浪费。

Proper training has content that is valuable and processes that can be implemented immediately. Training without implementation is a waste.

领导力。乍一看,与领导力相关的那些特质似乎与超级预测者的特质截然相反。调和这两者的方法是承认:恰当、甚至是大胆的行动,都需要良好的思考。而最优秀的领导者也明白,即便再周密的计划,也需要根据实际情况不断调整。

Leadership. At first blush, the qualities associated with leadership appear antithetical to those of the superforecasters. The way to reconcile the two is to acknowledge that proper, even bold, action requires good thinking. And the best leaders recognize that even the best laid plans need to be constantly revised based on the conditions.

附录:用布里尔评分法记分

Appendix: Keeping Score with Brier

心理学家通常使用布莱尔得分(Brier score)来衡量概率预测的准确性。

Psychologists commonly use the Brier score as a method for gauging the accuracy of probabilistic forecasts.

格伦·布里尔是一位气象学家,他在 1950 年代提出了这个评分方法。简而言之,布里尔评分衡量的是预测误差的平方,即(预测值 − 实际结果)²。对于二元事件,如果事件发生,实际结果取值为 1,否则为 0。就跟高尔夫球一样,分数越低越好。

Glenn Brier, a meteorologist, developed the score in the 1950s.30 In its simplest form, the Brier score measures the square of the forecast error, or (forecast − outcome)2. For binary events, the value of the outcome is 1 if the event occurs and 0 if it does not. As in golf, a lower score is better.

布里尔分数可以按 0 到 1 或 0 到 2 的标度来表述,具体取决于计算方法。我们遵循布里尔最初的思路,将结果置于 0 到 2 的标度上。按这种方式计算布里尔分数时,你要考虑事件发生和不发生这两种情况的预测误差平方。

You can express a Brier score either on a scale of 0 to 1, or 0 to 2, depending on the calculation. We follow Brier’s original approach and place our results on a scale of 0 to 2. When calculating the Brier score this way, you consider the squared forecast error for both the event and the non-event.

附录 5 展示了一位气象学家对接下来四天是否会下雨的概率预测。例如,在第 2 天,她预测下雨的概率为 80%。同样,我们可以说,她预测不下雨的概率为 20%。由于当天确实下雨了,我们在结果列中“下雨”下方填 1,在“不下雨”下方填 0。她那天的布里尔分数为 0.08。对于多次预测,总布里尔分数是每次预测得分的平均值。这位气象学家的总布里尔分数为 0.25。

Exhibit 5 shows a meteorologist’s probabilistic forecasts for whether it will rain over the next four days. For example, on Day 2, she forecasts an 80 percent probability that it will rain. Likewise, we can say she forecasts a 20 percent probability that it will not rain. Because it did rain, we place a 1 in the outcome column below “Rain” and a 0 in the “No Rain” column. Her Brier score for that day was 0.08. For multiple forecasts, the overall Brier score is the mean of the scores for each forecast. The meteorologist’s overall Brier score comes to 0.25.

表 5:布里尔分数的计算

Exhibit 5: Calculation of Brier Score

下雨无雨布里尔分数
预报结果预报结果计算结果
130%070%1= (0.3-0)² + (0.7-1)²0.18
280%120%0= (0.8-1)² + (0.2-0)²0.08
360%040%1= (0.6-0)² + (0.4-1)²0.72
4100%10%0= (1.0-1)² + (0.0-0)²0.00
均值0.25
   Rain   No Rain   Brier Score
Day Forecast Outcome Forecast Outcome   Calculation   Result
 1   30%   0   70%   1   = (0.3-0)2 +(0.7-1)2  0.18
 2   80%   1   20%   0   = (0.8-1)2 +(0.2-0)2   0.08
 3   60%   0   40%   1   = (0.6-0)2 +(0.4-1)2   0.72
 4   100%   1   0%   0   = (1.0-1)2 +(0.0-0)2   0.00
   Mean   0.25

资料来源:瑞信。

Source: Credit Suisse.

0 到 2 这个标度有一个不错的特性。随机猜测的布里尔分数精确等于 0.50。展图 6 展示了一个事件发生(“下雨”)时,主观概率从 0% 到 100% 所对应的布里尔分数。超级预测者的布里尔分数大约在 0.20 到 0.25 之间,在某些极个别情况下,可以达到十几的水平。

The scale from 0 to 2 has a nice feature. Random guesses have a Brier score of exactly 0.50. Exhibit 6 shows the Brier scores for an event that occurs (“Rain”) for subjective probabilities from 0 to 100 percent. Superforecasters have Brier scores of around 0.20 – 0.25 and in some exceptional cases can achieve scores in the teens.

展示图 6:不同主观概率下实际发生事件的布莱尔评分 2.0

Exhibit 6: Brier Scores of Event That Occurs for Various Subjective Probabilities 2.0

1.5

1.5

Brier Score 1.0

Brier Score 1.0

0.5

0.5

原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。

0.0 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 预测来源:瑞士信贷。

0.0 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 Forecast Source: Credit Suisse.

注释 1 菲利普·E·泰特洛克和丹·加德纳,《超预测:预测的艺术与科学》(纽约:皇冠出版社,2015 年),第 191 页。

Endnotes 1 Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (New York: Crown Publishers, 2015), 191.

2 菲利普·E. 泰特洛克,《专家的政治判断:它有多准?我们如何得知?》(新泽西州普林斯顿:普林斯顿大学出版社,2005 年)。

2 Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton, NJ: Princeton University Press, 2005).

3 戴维·伊格纳修斯,《闲聊过多》,《华盛顿邮报》,2013 年 11 月 1 日。

3 David Ignatius, “More Chatter Than Needed,” Washington Post, November 1, 2013.

4 丹尼尔·卡尼曼 与 加里·克莱因,《直觉型专业技能的成立条件:一场失败的歧见》,载于《美国心理学家》,第 64 卷,第 6 期,2009 年 9 月,第 515-526 页。

4 Daniel Kahneman and Gary Klein, “Conditions for Intuitive Expertise: A Failure to Disagree,” American Psychologist, Vol. 64, No. 6, September 2009, 515-526.

5 丹尼尔·卡尼曼,《思考,快与慢》(纽约:法勒、斯特劳斯和吉鲁出版社,2011 年)。

5 Daniel Kahneman, Thinking, Fast and Slow (New York: Farrar, Straus and Giroux, 2011).

6 Philip E. Tetlock,“良好判断项目:我们能改进对未来可能性的概率判断吗?”,在瑞士信贷思想领袖论坛上的演讲,2014 年 6 月 11 日。

6 Philip E. Tetlock, “The Good Judgment Project: Can We Improve Probability Judgments of Possible Futures?” Presentation at the Credit Suisse Thought Leader Forum, June 11, 2014.

要了解这一错误的更多内容,请参见菲尔·罗森茨维格所著的《光环效应……及其它八种误导管理者的商业错觉》(纽约:自由出版社,2007 年)。

7 To learn more of this mistake, see Phil Rosenzweig, The Halo Effect . . . and the Eight Other Business Delusions That Deceive Managers (New York: Free Press, 2007).

8 Kahneman, 212.

8 Kahneman, 212.

9 Keith E. Stanovich, 《智商测试遗漏了什么:理性思维的心理学》(纽黑文,康涅狄格州:耶鲁大学出版社,2009 年)。另见 Michael J. Mauboussin 与 Dan Callahan 合著,《IQ 与 RQ 之辩:

9 Keith E. Stanovich, What Intelligence Tests Miss: The Psychology of Rational Thought (New Haven, CT: Yale University Press, 2009). See also Michael J. Mauboussin and Dan Callahan, “IQ versus RQ:

“区分聪明与决策能力”,瑞信全球金融策略报告,2015 年 5 月 12 日。10 Jonathan Baron,“关于思维的信念”,载于 James F. Voss、David N. Perkins 和 Judith W. Segal 编,《非正式推理与教育》(新泽西州希尔斯代尔:Lawrence Erlbaum Associates, Inc.,1991 年),第 169-186 页。11 Charles Munger,“关于投资管理与商业的基本世俗智慧的一课”,《杰出投资者文摘》,第 10 卷,第 1 和第 2 期,1995 年 5 月 5 日,第 49-63 页。

Differentiating Smarts from Decision-Making Skills,” Credit Suisse Global Financial Strategies, May 12, 2015. 10 Jonathan Baron, “Beliefs about Thinking,” in James F. Voss, David N. Perkins, and Judith W. Segal, eds., Informal Reasoning And Education (Hillsdale, NJ: Lawrence Erlbaum Associates, Inc., 1991), 169-186. 11 Charles Munger, “A Lesson on Elementary Worldly Wisdom as It Relates to Investment Management & Business,” Outstanding Investor Digest, Vol. 10, No. 1 & 2, May 5, 1995, 49-63.

12 Michael J. Mauboussin, 《三思而后行:驾驭反直觉的力量》(波士顿,马萨诸塞州:哈佛商业出版社,2009 年)。

12 Michael J. Mauboussin, Think Twice: Harnessing the Power of Counterintuition (Boston, MA: Harvard Business Press, 2009).

13 詹姆斯·索罗维基,《群体的智慧:为何多数人比少数人更聪明,集体智慧如何塑造商业、经济、社会与国家》(纽约:道布尔迪出版社,2004 年)。

13 James Surowiecki, The Wisdom of Crowds: Why the Many Are Smarter Than the Few and How Collective Wisdom Shapes Business, Economies, Societies, and Nations (New York: Doubleday, 2004).

这与经济学家弗兰克·奈特提出的风险与不确定性之间的区别是一致的。塔德乌什·泰什卡和彼得·杰隆卡合著的《专家判断:金融分析师与天气预报员的对比》一文强调了专业判断在不同领域中可靠性的显著差异。

14 This is consistent with the difference between risk and uncertainty proposed by the economist Frank Knight. 15 Tadeusz Tyszka and Piotr Zielonka, “Expert Judgments: Financial Analysts Versus Weather Forecasters,”

《心理学与金融市场杂志》,2002 年第 3 卷第 3 期,第 152-160 页。

Journal of Psychology and Financial Markets, Vol. 3, No. 3, 2002, 152-160.

16 雷蒙德·S·尼克森,《确认偏误:一个以多种面目出现的普遍现象》,《普通心理学评论》,第 2 卷,第 2 期,1998 年 6 月,第 175-220 页。

16 Raymond S. Nickerson, “Confirmation Bias: A Ubiquitous Phenomenon in Many Guises,” Review of General Psychology, Vol. 2, No. 2, June 1998, 175-220.

17 Max Bazerman, 《管理决策中的判断》,第 4 版(纽约:John Wiley & Sons,1998)。 18 Carol S. Dweck, 《心态:成功的新心理学》(纽约:Random House,2006),第 18 页。

17 Max Bazerman, Judgment in Managerial Decision Making, 4th Ed. (New York: John Wiley & Sons, 1998). 18 Carol S. Dweck, Mindset: The New Psychology of Success (New York: Random House, 2006), 18.

19 安吉拉·达克沃斯,《坚毅:激情与毅力的力量》(纽约:斯克里布纳出版社,2016 年)。

19 Angela Duckworth, Grit: The Power of Passion and Perseverance (New York: Scribner, 2016).

20 迈克尔·J·莫布森(Michael J. Mauboussin)和丹·卡拉汉(Dan Callahan)合著,《打造高效团队:如何管理团队做出正确决策》(Building an Effective Team: How to Manage a Team to Make Good Decisions),瑞士信贷全球金融策略,2014 年 1 月 8 日。

20 Michael J. Mauboussin and Dan Callahan, “Building an Effective Team: How to Manage a Team to Make Good Decisions,” Credit Suisse Global Financial Strategies, January 8, 2014.

21 Elizabeth Mannix 和 Margaret A. Neale,“什么差异真正重要?组织内多元化团队的前景与现实”,《公共利益中的心理科学》,第 6 卷,第 2 期,2005 年 10 月,第 31-55 页。

21 Elizabeth Mannix and Margaret A. Neale, “What Differences Make a Difference? The Promise and Reality of Diverse Teams in Organizations,” Psychological Science in the Public Interest, Vol. 6, No. 2, October 2005, 31-55.

22 Jonathan Baron、Barbara A. Mellers、Philip E. Tetlock、Eric Stone 和 Lyle H. Ungar,《让聚合概率预测更加极端的两个理由》,《决策分析》第 11 卷第 2 期,2014 年 6 月,第 133–145 页。

22 Jonathan Baron, Barbara A. Mellers, Philip E. Tetlock, Eric Stone, and Lyle H. Ungar, “Two Reasons to Make Aggregated Probability Forecasts More Extreme,” Decision Analysis, Vol. 11, No. 2, June 2014, 133- 145.

Garold Stasser 与 William Titus,“群体决策中未共享信息的汇集:讨论过程中的信息取样偏差”,Journal of Personality and Social Psychology,第 48 卷,第 6 期,1985 年 6 月,第 1467–1478 页。另见 Jennifer R. Winquist 与 James R. Larson, Jr.,“信息汇集:何时影响群体决策”,Journal of Personality and Social Psychology,第 74 卷,第 2 期,1998 年 2 月,第 371–377 页。

23 Garold Stasser and William Titus, “Pooling of Unshared Information in Group Decision Making: Biased Information Sampling During Discussion,” Journal of Personality and Social Psychology, Vol. 48, No. 6, June 1985, 1467-1478. Also, Jennifer R. Winquist and James R. Larson, Jr., “Information Pooling: When It Impacts Group Decision Making,” Journal of Personality and Social Psychology, Vol. 74, No. 2, February 1998, 371-377.

芭芭拉·梅勒斯、埃里克·斯通、特里·默里、安吉拉·明斯特、尼克·罗尔巴赫、迈克尔·毕晓普、陈伊娃、约书亚·贝克、侯媛、迈克尔·霍洛维茨、莱尔·昂加尔和菲利普·泰特洛克,《识别与培养》

24 Barbara Mellers, Eric Stone, Terry Murray, Angela Minster, Nick Rohrbaugh, Michael Bishop, Eva Chen, Joshua Baker, Yuan Hou, Michael Horowitz, Lyle Ungar, and Philip Tetlock, “Identifying and Cultivating

“超级预测者:一种改进概率预测的方法,”《心理科学展望》,第 10 卷,第 3 期,2015 年 5 月,第 267–281 页。

Superforecasters as a Method of Improving Probabilistic Predictions,” Perspectives on Psychological Science, Vol. 10, No. 3, May 2015, 267–281.

25 Caleb Melby、Laura Marcinek 和 Dani Burger,《美联储批评者称 2010 年通胀警告信仍然正确》,彭博社,2014 年 10 月 2 日。

25 Caleb Melby, Laura Marcinek, and Dani Burger, “Fed Critics Say ’10 Letter Warning Inflation Still Right,” Bloomberg, October 2, 2014.

26 Tetlock (2014).

26 Tetlock (2014).

27 参见 https://en.wikipedia.org/wiki/Fermi_problem。

27 See https://en.wikipedia.org/wiki/Fermi_problem.

威廉·H·德沃斯奇斯、F·里德·约翰逊、理查德·W·邓福德、凯文·J·博伊尔、萨拉·P·哈德森以及 K.

28 William H. Desvousges, F. Reed Johnson, Richard W. Dunford, Kevin J. Boyle, Sara P. Hudson, and K.

妮可·威尔逊,《使用条件价值法衡量非使用损害:准确性的实验评估》,第二版(北卡罗来纳州三角研究园:RTI 出版社,2010 年)。

Nicole Wilson, Measuring Nonuse Damages Using Contingent Valuation: An Experimental Evaluation of Accuracy, 2nd Ed. (Research Triangle Park, NC: RTI Press Publication, 2010).

29 纳西姆·尼古拉斯·塔勒布,《黑天鹅:如何应对不可预知的未来》(纽约:兰登书屋,2007 年),第 272 页。

29 Nassim Nicholas Taleb, The Black Swan: The Impact of the Highly Improbable (New York: Random House, 2007), 272.

格伦·W·布赖尔,《以概率形式表述预测的验证》,《每月天气评论》,第 78 卷第 1 期,1950 年 1 月,第 1–3 页。

30 Glenn W. Brier, “Verification of Forecasts Expressed in Terms of Probability,” Monthly Weather Review, Vol. 78, No. 1, January 1950, 1-3.