Kesselman List of Estimative Words (fig 5.2) (gwern finds this usefull for LLMs and so do I). These are probability words calibrated to meaning where the general public understand them to be, with ambiouso words avboided.
Words of estimative probability (WEP or WEPs) are terms used by intelligence analysts in the production of analytic reports to convey the likelihood of a future event occurring.
| Word | Certainty |
|---|---|
| Almost Certain | 86-99% |
| Highly Likely | 71-85% |
| Likely | 56-70% |
| Chances a little better [or less] than even | 46-55% |
| Unlikely | 31-45% |
| Highly Unlikely | 16-30% |
| Remote | 1-15% |
This issue prompts the researcher to suggest the Kesselman List of Estimative Words
for use within the intelligence community. It builds on Sherman Kent’s original WEP list in
the 1960s and the National Intelligence Council’s current list as well as draws from
Mercyhurst College’s WEP list. The new scale includes seven words of estimative
probability which is in line with what Kent and the NIC have proposed; however, it differs in
its phraseology and odds equivalents. The percentile ranges are broken down into groups of
15%, except for the middle range of chances a little better [or less] which was assigned only
10% and the upper and lower ranges which number 14%. Absolute certainty or impossibility
generally is not conveyed in intelligence assessments, but the two extremes are represented at
the top and bottom of the new scale.
Most importantly, the list uses words that large groups of people perceive in similar
manners. Subjects have never had problems with the extreme ends of a scale; therefore,
perceptions of almost certain and remote as well as highly likely and highly unlikely should
remain fairly constant. It is important to note that these terms are present in the NIC’s new
word list, although they vary slightly. Rather than using almost certainly, this researcher
believes that it is possible to use almost certain in far more grammatical structures, thus the
elimination of –ly. Second, the word highly conveys a much clearer picture with likely and
unlikely than does very, so those two phrases were also tweaked.
Where the problems appear, however, are with the words buried in the middle of the
scale. In weather forecasting, respondents of the Juneau survey indicated that the word likely
conveyed a 62.5% numerical equivalent and in the medical professions physicians have
indicated that the word’s value is approximately 70%. An odds equivalent in the scale above
of 56-70% mirrors that of researchers’ findings in several disciplines. Most notably, use of
only the word likely, rather than likely/probably as synonyms for one another (as in the NIC’s
new scale), should serve to eliminate confusion and standardize that particular percentile range.
The next question to tackle was how to convey odds that fell directly above or below
50%. The terms chances are even and fifty-fifty tells a decision maker nothing and
essentially asks them to toss a coin in the air. Therefore, a term was needed that would
convey odds slightly above or below the halfway benchmark and chances a little better [or
less] does exactly this. Only assigning the category 10% forces the analyst to make a call
depending on whether the chances are indeed better or less than a particular situation coming
to fruition. For example, if an analyst wants to convey that certain odds are better, they would have to equate the statement with at least 51%. It may sound as if 51% is not much different than 50%, but saying in essence that there is the slightest probability something may occur is a progress within the IC.
Finally, this list is not only extremely easy to use but also simple for analysts to
remember and produce on their own if they did not have a copy with them. The words are
arranged in such a way that the top of the list generally mirrors that of the bottom (with the
exceptions of almost certain and remote): highly likely realizes its counterpart in highly
unlikely and likely and unlikely mirror each other as well. Analysts simply need to remember
that each category is broken down into groups of 15% except for the middle category at 10%
and the two upper boundaries which will never reach complete certainty or impossibility.
The scale also eliminates the need for synonyms in estimative language. If these words can
be used consistently, analysts will always know exactly what they are attempting to convey
and decision makers will receive clarity, allowing them to enact policy that is in line with an
analyst’s thinking.
While there is an understanding of the value of consistent terminology in the IC, it
has yet to be operationalized. With the increased scrutiny that the IC is likely to receive in
the new information age, it is only to their benefit to adopt such a list. While the Kesselman
List of Estimative Words will likely be tweaked by others in the community, it is a step in the
right direction. Analysts have an obligation to communicate as effectively as they can the
results of their estimates. The best case scenario is that the NIC and the IC take into
consideration this new estimative scale above and produce several more iterations of their
own list until employees of the community can come to agreement on a set of clear-cut words
that all are both willing to accept and employ in daily practice.
From fig 5.2 in Kesselman's 2008 List of Estimative Words. See also Kent's words of estimative probability, 1964