85 drawings, then 60, then a multiple-choice backup
The test began as part of the authors' aphasia work at the Boston Veterans Administration Hospital and Boston University. It drew on earlier naming tests, especially the Oldfield Object Naming Test of 1971, which weighed word frequency, familiarity and visual complexity.8 The 60 drawings were chosen from the 85 to span a range of difficulty.5
Same answers, different scores
Different scoring approaches can produce large differences in total scores.1 Below is a fictional record of 30 items. Tap any item to change how it was answered, then switch the scoring rule and the stopping rule.
A fictional 30-item record, not BNT items or norms. The real test has 60 items, and its own rules on where to start aren't modelled here.
"Houston, we have a problem"
That was the title of a survey of how clinicians actually used the test. Answers varied on which version, which norms and which rules.1
Where it's used
References
- [1]Survey results regarding use of the Boston Naming Test: Houston, we have a problem. ResearchGate.researchgate.net
- [2]The Boston Naming Test: calculation and utilization of an estimated 60-item score from 30- and 15-item administrations. ResearchGate.researchgate.net
- [3]Olabarrieta-Landa, L., et al. (2015). Standard form of the Boston Naming Test: normative data for the Latin American Spanish speaking adult population. NeuroRehabilitation. IOS Press.content.iospress.com
- [4]Boston Naming Test. In Encyclopedia of Clinical Neuropsychology. Springer reference entry (institutional mirror).link.springer.com (mirror)
- [5]An empirically derived short form of the Boston Naming Test. Archives of Clinical Neuropsychology. ScienceDirect.sciencedirect.com
- [6]Boston Naming Test scoring and documentation. BastionGPT (2026).bastiongpt.com
- [7]Transcranial magnetic stimulation pilot in primary progressive aphasia: protocol NCT05860647. ClinicalTrials.gov (PDF).cdn.clinicaltrials.gov
- [8]Boston Naming Test. Grokipedia (2026).grokipedia.com
Related historical tests
The PPVT says a word and asks for the matching picture; the BNT shows a picture and asks for the word. One tests understanding, the other retrieval.
Naming scores are usually read against broader ability. A full adult battery like the WAIS gives that context.
Like the BNT, the d2 targets one narrow ability. Both are used in clinical trials as focused outcome measures.
The WAIS-R showed how much age and cohort shape scores. Naming scores also vary with age, education and gender, which is why BNT norms matter so much.
Hamilton's scale shows how the way items are summed can mislead. The BNT survey found the same problem with inconsistent scoring rules.