Four requirements, set before any item was written
Otis opened the paper with a specification. He wrote that "the chief object of testing in groups, of course, is economy of time", and that the responses should therefore be "very simple". Most of the test ideas, he said, came from the Stanford Revision of the Binet Scale [1].
Ten tests, each written twice
The scale was "compiled in duplicate": two versions of every test, as alike as possible without repeating material, so that each pupil's two scores could be compared to measure reliability. 121 pupils from grades 4, 6 and 8 of one large grammar school took both versions [1]. Select a test to see its format.
Rights minus wrongs
Where an item had only two possible answers, Otis reasoned that a pupil guessing at unknown items would get half of them right by luck. So he subtracted one point for every wrong mark: blanks cost nothing, and someone who guessed every item would expect a score of zero [1]. Try both two-choice tests, then compare the raw count with Otis's score.
Points, not years, and a quotient called brightness
Each raw score was converted to points so that every test carried equal weight, and the ten point scores were added. The age norms for 10 to 14 fell almost on a straight line, about 40.5 points a year. Extended backwards, that line reached zero points at age zero, which Otis called "an entirely accidental coincidence". It meant points could be divided directly: total points over the norm for the pupil's age gave a Coefficient of Brightness[1].
480 points ÷ 445 (the smoothed norm at 11.0) = 1.08. Otis trusted only the norms for ages 10 to 14, and called those above 15 "only roughly approximate", because older pupils still in grammar school were a selected group held back in school.
The grades overlapped
Average point scores were 385 in fourth grade, 514 in sixth and 605 in eighth. Yet 8 of the 43 fourth graders scored above Otis's fifth-grade norm, and one above the sixth. Within a single grade the youngest pupils tended to score highest. Otis read that as a sign that bright children were being held back by an "inelastic system of grading" [1].
From a journal study to a publisher's catalogue
Otis signed the 1918 paper from the Surgeon General's Office in Washington, which is where his wartime testing work had taken him. The scale he built at Stanford fed the Army's group examination; after the war he turned it into a product, then worked for decades as a test editor.
Corrections to the common story
Earlier versions of this page, and much of what is written about Otis online, repeated claims that the sources don't support. Here is each one, checked.
References
- [1]Otis, A. S. (1918). An absolute point scale for the group measurement of intelligence. Journal of Educational Psychology, 9(5–6). doi:10.1037/h0072885. Reprint, University of Toronto copy, Internet Archive.archive.org
- [2]Otis, A. S. (1918). An absolute point scale for the group measurements of intelligence, Part I. Record with DOI. Zenodo.zenodo.org
- [3]Psychological Test, Otis Group Intelligence Scale (MA.316371.050). National Museum of American History, Smithsonian Institution.americanhistory.si.edu
- [4]Psychological Test, Otis Self-Administering Tests of Mental Ability, Higher Examination: Form A. National Museum of American History.americanhistory.si.edu
- [5]Leaflet with Sheet, Otis Quick-Scoring Mental Ability Tests: Manual of Directions for Gamma Test. National Museum of American History (citing Siegel 1992 and Gade 2018).americanhistory.si.edu
- [6]Arthur Sinton Otis: background, education and career. Prabook.prabook.com
- [7]Boake, C. (2002). From the Binet–Simon to the Wechsler–Bellevue: tracing the history of intelligence testing. Journal of Clinical and Experimental Neuropsychology, 24(3). Academia.edu copy.academia.edu
- [8]History of the Otis-Lennon School Ability Test. K12 Academics.k12academics.com
- [9]Otis Group Intelligence Scale (1929 set). Randolph College Psychology Collection, eHive.ehive.com
- [10]Test Collection Database: Otis. Social Sciences, Health and Education Library, University of Illinois.library.illinois.edu
- [11]Arthur Otis (1886–1964). Psychology Encyclopedia, JRank.psychology.jrank.org
Related historical tests
Some of Otis's written tests went into the Alpha battery, and he helped analyse its 1917 trials. The Alpha's disarranged sentences and analogies follow the format of his 1918 scale.
Otis wrote that his test ideas came chiefly from Terman's Stanford Revision. He also checked his scale against it, finding a mental-age error about the same as the Binet's.
Trabue's completion test was one of the fifteen tests in Otis's preliminary study of 29 pupils. Its sentence-completion format is close to Otis's own Narrative Completion test.
Otis's 1936 redesign took his aim of quick, objective scoring further. It came in Alpha, Beta and Gamma levels, published by World Book.
Roger Lennon's revision carried Otis's name into the modern era. It is the line through which the 1918 scale's group-testing approach survives in schools.
Pyle's battery came out five years before Otis's scale and was also meant for testing schoolchildren. Otis added what Pyle's battery lacked: equal-weight point scores, and a reliability figure for every test.