A Tape Dictionary for Linguistic Experiments

A TAPE DICTIONARY FOR LINGUISTIC EXPERIMENTS*
J. L. Dolby, H. L. Resniko.(f, and E. MacMurray * *
INTRODUCTION
phonetic forms consisting of an updated version of the 20,000 most used words from the
Thorndike Senior Century Dictionary. Other
efforts are known to be underway at Indiana
University2 and Standford University. Further,
a number of study groups have prepared small
dictionaries containing syntactic information
(e.g., the work at Harvard 3 ) but these have
generally been set up to handle specific texts
and expanded only as the need arises.
Almost since the arrival of the first piece of
modern computing equipment there has been
considerable interest in the potential of this
equipment for various aspects of linguistic data
handling. Experiments in machine translation
date to at least the late forties and more recently experiments in indexing and abstracting
are becoming more widespread. Almost without
exception these experiments have shown that
the basic problems are not trivial and that
serious experimentation on an extensive level
will be necessary before any of these problems
can be conquered-even assuming that solutions of some sort can be found. As a result
there has been an increasing amount of speculation about the fundamental nature of linguistic
structure and a growing need for fundamental
data on which various conjectures about this
structure can be tested.
Early experiments at Lockheed on automatic
parsing (Earl, Ricklefs, and Robison 4 ) demonstrated that much progress could be made
in this area and that a complete dictionary or
some combination of a dictionary and a part-ofspeech algorithm would be necessary as input
to the parsing program. As a result, it was
decided to key punch the words of the Pocket
Oxford Dictionary together with the parts of
speech given there in order to determine the
extent to which a part-of-speech dictionary
could be replaced by a small dictionary and a
part-of-speech algorithm. It was found (Resnikoff and Dolby5) that it was possible to
construct an algorithm that would rival the
accuracy of the Pocket Oxford when used in
conjunction with a small (200 word) dictionary. However, it was also found that neither
the algorithm nor the Pocket Oxford provided
sufficiently accurate and detailed information
for the parsing studies.
Recently several word lists and dictionaries
have come into being to help meet this need.
The University of Pennsylvaniachas compiled
a large word list consisting of the union of the
non-obsolete words in the Second Edition of
Webster's New International Dictionary and
four technical dictionaries. As with most of
the lists mentioned here, the Pennsylvania list
is available in alphabetic order and reverse
alphabetic order (that is, the words are arranged alphabetically beginning with the last
rather than the first letter). Cornell University
has produced a tape dictionary on written and
Concurrently, research on the nature of word
breaking, or graphemic syllabification (Dolby
*This work is supported by the Lockheed Independent Research Program.
**Mr. Dolby is with the Lockheed Missiles and Space Company, 3251 Hanover St., Palo Alto, California. Mr.
Resnikoff is presently at the University of Munich, Germany, on leave from Lockheed and Mr. MacMurray is with
the Statistical Tabulating Corporation. San Francisco.
419
From the collection of the Computer History Museum (www.computerhistory.org)
420
PROCEEDINGS-FALL JOINT COMPUTER CONFERENCE, 1963
and Resnikoff 6 ), showed that it would be
highly desirable to construct an extensive list
of words showing the legitimate break positions
(by dictionary standards) together with major
and minor stress positions. Further, when it
became evident that graphemic syllabification
was intimately connected with the algorithmic
determination of parts of speech (Dolby and
Resnikoff7) it was obvious that these two types
of information should be combined into a single
dictionary available for further computer experiments.
In this paper we discuss the sources of information used to construct such a dictionary,
the problems in attaining a reasonable degree of
accuracy and consistency in such a task and
the various output forms that have proven to
be useful in this research.
Dictionary Sources
As a first step in the construction of this
dictionary, all of the one-syllable words of the
Oxford Universal Dictionary were studied in
considerable detai1. 7 It was found that the
dictionary part-of-speech designations for
these words varied from one source to another
but that when all sources were considered the
designations were quite consistent. By "all"
in this case, we mean the parts of speech given
in both the 2nd and 3rd editions of Webster's
New International, the Oxford Universal Dictionary and the Oxford English Dictionary. It
was also found that the consistency increased
markedly as the words considered were narrowed down to what we might term the standard words of the language. Thus it was evident
that it would be necessary to merge the information from at least two sources to obtain
a measure of dictionary consistency and that
some information as to the relative utility of
the word would also be needed.
As a first approximation, it was decided to
obtain all of the left-justified, bold-faced words
of the Shorter Oxford together with the partsof-speech given by that dictionary and to add
to this the parts of speech for these words
found in the 3rd edition of Webster's New International Dictionary. In addition, a code was
supplied for each part of speech given in each
source to designate the status given. If any
standard meaning was to be found in that
source for that part of speech the word was
called standard for that part of speech. Otherwise, the first non-standard designation (such
as obsolete, dialectical, archaic, etc.) was used.
Figure 1 illustrates a typical page of material
as entered by the keypunch operators. Table 1
provides a list of codes used.
TABLE 1
Status Codes
Dictionary
Form
Dialectical
Alien
Archaic
Colloquial
Capital
Erroneous
Nonsense
Nonce Word
Obsolete
Poetical
Rare
Rhetoric
Specialized
Standard
Substandard
Keypunch
Code
Edited
Code
DI
AL
AR
CO
CA
ER
NS
NW
OB
PO
RA
RH
SP
ST
SS
D
F
A
Q
C
E
N
W
0
P
R
H
$
S
Z
Part-of-Speech Codes
1 Noun
2 Adjective
3 Verb
4 Adverb
5 Preposition
6 Conjunction
7 Pronoun
8 Interjection
9 Past
0 Other
The second stage of the operation consisted of
a relatively simple editing operation designed
to put the information in more usable form.
This consisted of several steps. First, the coding scheme originally adopted was designed
primarily to simplify the keypunching operation in the interest of minimizing keypunch
errors. As a result, a three letter code was
used for each part-of-speech, status entry. The
overall format was such as to restrict the total
From the collection of the Computer History Museum (www.computerhistory.org)
A TAPE DICTIONARY FOR LINGUISTIC EXPERIMENTS
?<I.ain entry
CAUSE
CAUSE CELEBRE
CAUSELESS
CAUSERIE
CAUSEUSE
CAUSEWAY
CAUSEY
CAUSIDICAL
CAUSON
CAUSTIC
CAUSTICITY
CAUTEL
CAUTER
CAUTERANT
CAUTEfiISM
CAUTERIZE
CAUTERY
CAUTION
CAUTIONARY
CAUTIOUS
CAVA
CAVALCADE
CAVAL! ER
CAVALLY
CAVALRY
CAVATINA
CAVE
CAVEAT
CAVEL
CAVENDISH
CAVERN
CAVERNOUS
Part-of-speech and Status Codes
Webster I s Third
Shorter Oxford
lST3ST6ST
1ST
2ST
1ST
1ST
lST3ST
lST3DI
lST3ST6DI
lAL
2ST
lAL
lAL
1ST
lST3DI
2ST
lOB
lSP2ST
1ST
lOB
1ST
1ST
lOB
3ST
1ST
lST3ST
lST25T
2ST
lAL
lST3ST
lST2ST
1ST
lST2ST
lAL
lST20B3ST8AL
lST30B
1DI3DI
1ST
lST3ST
25T
lST2ST
1ST
lOB
1ST
lST2ST
3ST
1ST
lST3ST
lST2ST
2ST
lSTOST
lST3ST
lST25T35T
lST2ST
1ST
lST2ST3ST
lST3ST
1DI3DI
1ST
15T2ST3ST
2ST
421
Sequence
Number
1107020
1107030
1107040
1107050
1107060
1107070
1107080
1107090
1107100
1107110
1107120
1107130
1107140
1107150
1107160
1107170
1107180
1107190
1107200
1107210
1107220
1107230
1107240
1107250
1107260
1107270
1107280
1107290
1107300
1107310
1107320
1107330
Figure 1
number of such entries for one word in one
dictionary to seven. A few words (such as a)
had more than seven parts of speech. These
added entries were provided on a second card.
In the editing run these card-pairs were reduced
to a single record on the tape. At the same
time, the part-of-speech, status field was reduced to ten columns for each dictionary by
representing the part-of-speech by column position and by using a single literal code for the
status. Further, the word field was reduced to
twenty columns since inspection of the original
printout showed that only five words of the entire 73,000 exceeded that number of letters.
In addition, it was possible to add certain
derived information to the record. A second
word field of twenty columns was added to store
the reversed form of the word in left justified
position to enable us to obtain "backwards"
orderings with standard sorting routines. The
backwards ordering is useful in a number of
ways, the most important being that it permits
us to obtain all the words with the same ending
(e.g., the bame suffix). Another column was
devoted to the coding of the status of the word
(as opposed to the status of its meanings given
in the part-of-speech fields;. The code adopted
was b if there was at least one standard partof-speech code in both of the dictionary partof-speech fields, x if only the Shorter Oxford
provided a standard meaning for the word, W
if only Webster's provided a standard meaning
for the word and a blank if neither source provided such a meaning.
Yet another column was devoted to recording a numerical code to approximate the number of syllables in the word since, as already
. noted, there is now known to be an intimate
connection between the number of syllables in
the word and its part of speech assignments.
The approximation to this was obtained by
counting the number of vowel strings in the
word entry (but not counting a final e as a
vowel). The same column was also used to
provide an overriding code for prefixes, suffixes, hyhenated words and broken words (the
codes being p, s, hand b respectively). This
latter determination was made by quite obvious
use of the presence of hyphens at the beginning, end or intermediary portions of the words
and by the presence of an intermediary blank
in the word.
Finally, a new field of ten columns was devoted to obtaining a merged part-of-speech,
From the collection of the Computer History Museum (www.computerhistory.org)
422
PROCEEDINGS-FALL JOINT COMPUTER CONFERENCE, 1963
§a
u
v
...-I
~§
...-I.,
Main Entry
Part-of-speech &: Status
...-I'"
~~ Shorter
Reversed Entry
Sequence
Oxford
CAUSE
CAUSE CELEBRE
CAUSELESS
CAUSERIE
CAUSEUSE
CAUSEWAY
CAUSEY
CAU5IDICAL
CAUSON
CAUSTIC
CAUSTICITY
CAUTEL
CAUTER
CAUTERANT
CAUTERISM
CAUTERIZE
CAUTERY
CAUTION
CAUTIONARY
CAUTIOUS
CAVA
CAVALCADE
CAVALIER
CAVALLY
CAVALRY
CAVATINA
CAVE
CAVEAT
CAVEL
CAVENDISH
CAVERN
CAVERNOUS
ESUAC
ERBELEC ESUAC
SSELESUAC
EIRESUAC
ESUESUAC
YAWESUAC
YESUAC
LACIDISUAC
IB5 5
BWF
3B 5
4WF
3WF
4B5
2BS D
4X 5
2 0
2B$S
4B5
2 0
2B5
3B5
3 0,
3B S
3BS
2B5 S
4BS5
4B S
2WF
3BS S
3BSS
3XS
3B5S
4WF
1BSOS
2BS 0
2 D D
3B5
2B5 5
3B S
~OSUAC
C ITSUAC
YTICITSUAC
LETUAC
RETUAC
TNARETUAC
MSIRETUAC
EZ IRETUAC
YRETUAC
NOITUAC
YRANOI TUAC
suor TUAC
AVAC
EDACLAVAC
REILAVAC
YLLAVAC
YRLAVAC
ANITAVAC
EVAC
TAEVAC
LEVAC
HSIDNEVAC
NREVAC
SUONREVAC
D
Websters
5 5
S
S
S
5
S S
5 D
55
S
0
S
5S
~
S
S S
SS
S
5
S S
SSS
F
SS
S
SSS
S S
D D
S
S5S
S
5
#
Merged
S 5
S
5
S
S
5 S
5 D
S
0
55
5
0
S
5S
0
S
S
S S
SS
5
S5
S S
55S
S
S5
S
5S5
S 5
D D
S
SS5
S
5
1107020
1107030
1107040
1107050
1107060
1107070
1107080
1107090
1107100
1107110
1107120
1107130
1107140
1107150
1107160
1107170
1107180
1107190
1107200
1107210
51107220
1107230
11('172';'0
1107250
1107260
1107270
F ),107280
Il07290
1107300
1107310
1107320
1107330
Figure 2
status field derived from the information given
by each dictionary. This again was accomplished in a rather obvious manner giving precedence to standard meanings or the first meaning given (in this case, the one given by the
Shorter Oxford) whenever the two sources did
not agree.
The resulting tape record was thus again
expanded to a full eighty columns as is shown
in the sample page given in Figure 2. However,
in this form it is now possible to sort the information in a number of ways by standard
sort routines. Generally, the most useful sort
pattern is to provide a major sort on the "number of vowel string" code followed by a minor
sort on either forward alphabetic or reverse
alphabetic order.
Additions to the Original Source Material
Our original choice of the Shorter Oxford as
a word list was made because it provided a list
of manageable size with excellent part-ofspeech information. However, there remained
some question in our minds as to the extent of
coverage of this source for common American
usage. Fortunately, R. L. Venezky was kind
enough to make available to us the Cornell
University tape of 20,000 commonly used
words. All of the words in this list that were
not contained in the Shorter Oxford were
added to original source together with the partof-speech information for these words provided by TVebster's 3rd. There were 2490
of these words, incidentally.
The next main step in this area is still in
progress. As we have already noted, graphemic
,syllabification is of considerable interest to us.
We have therefore undertaken to punch all of
the words of Funk and Wagnall's New Practical Standard Dictionary in the syllabified
form given by that source together with the
accent information there given. The choice of
source for this information stemmed from the
fact that this was one of the few dictionaries
to provide this information on the derived
forms in a reasonable manner for direct keypunching. When this list is complete, the
syllabification information will be added to the
existing information for the words in the
original source. (It is not contemplated at
this time that all of the words in the Funk and
Wagnall's list will be added to the list since it
From the collection of the Computer History Museum (www.computerhistory.org)
A TAPE DICTIONARY FOR LINGUISTIC EXPERIMENTS
now appears that some 50,000 added entries
would be necessary.) When this is done, the
present "vowel-string approximation" will be
replaced by the actual Funk and Wagnall's
syllable count. (It will, however, be of interest
to study precisely where the two differ.) Another check of some interest will be the comparison of the Thorndike usage factors given
in the Cornell list with the rather crude status
designation we have derived on the basis of
standard word designations given by our two
sources of part-of-speech information.
Accuracy and Consistency
A brief study of the structure of almost any
dictionary is sufficient to show that dictionaries
are constructed primarily for human, as opposed to machine usage. As a result, a number
of conventions had to be established in order
to provide a reasonable degree of consistency
in the coding. For example, some five different
forms were used by the Shorter Oxford to indicate obsolete words (with a number of
variations within one of these forms). Further,
we found it very difficult to determine in advance a precise set of rules that would resolve
all the questions that were to arise. As a result, a rough structure of the coding was laid
out in advance and this was filled in, in detail,
as problems came up. A careful log of all decisions was maintained throu,ghout so that
once a form was encountered it could be
checked against the log to see if a decision had
been made as to how it should be coded.
The generation of the main list was verified
completely with additional spot checks throughout the project. The key punching was also
verified one hundred percent. A random sample
of 280 words was then chosen from the entire
list and these were checked and no errors were
found. In addition, the symmetric difference
between the main list and the Cornell list was
checked closely to determine whether any of the
words in either difference set could be there due
to an error in either list.
Users of the list have been kind enough to
report errors as they are found. The spelling
research group at Stanford, for instance, has
made an intensive study of a random sample
of 1500 words from the main list and they re-
423
port that no spelling errors were found in that
sample.
SUMMARY
A tape dictionary of sOlne 75,000 entries has
been prepared with part-of-speech, status,
usage, graphemic syllabification and stress information. The entries have been sorted
alphabetically forward and backward as well
as by syllable and by part-of-speech. Comparisons are being drawn between various measures
of usage as well as between two measures of
the number of syllables in the written form.
Considerable care has been taken to minimize
the number of errors in the list and to insure
a high degree of consistency in the coding. The
authors believe that the resulting listing will be
of great utility in basic studies of the nature
of linguistic data handling.
REFERENCES
1. Normal and Reverse English Word List
(8 Volumes). Compiled under the direction of A. F. BROWN, University of
Pennsylvania (1963).
2. ASTIA Report AD 273 500. Seventh
Quarterly Report on Automatic Language
Analysis.
3. Mathematical Linguistics and Automatic
Translation. Report No. NSF-8. Harvard
University, Cambridge, Massachusetts.
January 1963.
4. "Automatic Syntactic Analysis of Simple
English Sentences." L. L. EARL, B. RICKLEFS, H. R. ROBISON. Lockheed Document
6-90-61-40. July 31, 1961.
5. RESNIKOFF, H. L., and DOLBY, J. L., Correspondence of the July 1963 Proceedings
of the I.E.E.E.
6. "On the Structure of Graphemic Syllabification," J. L. DOLBY and H. L. RESNIKOFF. Presented at the August 1963 Meeting of the Association for Machine Translation and Computationat Linguistics,
Denver. (Abstract appears in Mechanical
Translation, August 1963.)
7. "Prolegomena to a Study of Written
English," J. L. DOLBY and H. L. RESINKOFF. Lockheed Document 6-90-63-5.
February 1963.
From the collection of the Computer History Museum (www.computerhistory.org)
From the collection of the Computer History Museum (www.computerhistory.org)