The LEADER retrieval system

The LEADER retrieval system
by DOXALD J.
HILL~IAX
and ANDREW J. KASARDA
Lehigh University
Bethlehem, Pennsylvania
INTRODUCTION
The LEADER system is a new service-oriented prototype designed to meet the retrieval needs of research
scientists· working within or in conjunction with the
Center for the Information Sciences at Lehigh University. In the first part of this paper, we describe the
major conceptual apparatus and principal design
features· of LEADER, while the second part contains
a brief discussion of system implementation and user
interaction.
The name "LEADER" is an acronym for "LEhigh
Automatic Device for Efficient Retrieval," and is thus
similar to other acronyms in possessing both an intended
meaning as well as an actual referent. Imaginative
readers can undoubtedly supply alternative and presumably more ribald interpretations of the same six
characters, but this is rather incidental to the main
goal of the LEADER system, which is to provide a
very highly user-oriented .facility for .the negotiation of
open-ended inquiries and interactive browsing. To
help meet this objective, the system includes on-line
processing of requests, using a novel and relatively
inexpensive hardware configuration, and serially organizes its output in the form of document references,
dtations to documents, and complete textual passages
selected from one or several documents, in any way that
the user specifies. This ability of the user to control output is but one feature of an overall interactive procedure
which begins when an initial request is entered into the
LEADER system in the form of a set of sentences
describing the user's problem. Each input sentence
must, of course, be grammatically well-formed, but
there is no restriction on vocabulary. A typical inquiry
might read:
"I would like to know whether modular bounded
functionals have ever been used in theoretical
studies of retrievable sets, and if so by whom and
with what results. If there has been no application
of this type, I would be interested to learn of any
work in retrieval theory that makes use of Borel
functions. If there is no such work, please direct
me to retrieval studies involving topological measures or metric spaces in generaL"
Inquiries such as these are presented directly to the
system and displayed on a C~T scope. As each inquiry
is displayed, it is also automatically analyzed by the
same procedures used to process the full text of input
documents. That is to say, LEADER treats both documents and queries as entities of the same logical type to
begin with, so that the logical and referential structure of an inquiry is accorded just as much importance
as the structure of a document. The goal of text processing is therefore the same throughout, viz., to determine what each group of input sentences is about,
whether they constitute a document or an inquiry, and
to establish major patterns of conceptual relatedness
between documents and terms used either in document
or query characterization. The text-processing .features
of LEADER thus include elements of syntax, semantics,
and logic.
After the sentences of an initial inquiry have been
analyzed into concept-denoting expressions and their
logical interrelationships, LEADER is able to fashion
an appropriate response to the user's retrieval needs by
comparing the conceptual structure of the inquiry with
the general structure of the data base. This comparison
is conducted via a man/machine dialogue in which
LEADER instructs and interrogates the user, attempting to acquaint him with the nature of its stored information so that each inquiry can be negotiated
through successive modifications of the user's stated
interests. The dialogue itself is carried out on a CRT
scope.
The user may call for document references, citations,
or passages of relevant text at any time during the
negotiation, so that by a process of selective browsing
447
From the collection of the Computer History Museum (www.computerhistory.org)
448
Spring Joint Computer Conference, 1969
he may assist LEADER to arrive at the most appropriate solution to his retrieval problem. What the user
wants in the way of final output is a matter for him to
decide. In most cases, uSers prefer to read portions of
the text of seiected documents on the CRT scope, and
to ask for hard copy of what they state to be the most
interesting or pertinent of the documents that have
been displayed for them on the CRT. Such hard copy
is provided by either a line printer or a low-speed terminal.
Work on the theoretical basis of LEADER was begun in '1962, and preliminary implementation of the
theory took place in 1965-1966, when programs for fulltext analysis were written and tested. In 1966-67, system components for connectivity and decision were
developed, and a small scale prototype went into operation in the SlLTfiffier of 1967. A full scale version of
LEADER is now virtually complete, employing a
data base of technical articles and reports in information science, the full text of which is in machine-readable
form.
The design features of the system call for a substantially different approach to evaluation, largely
because such widely-used and familiar notions as "precision" and "recall" have no operational meaning for
LEADER. For this reason, we have proposed new
methods for evaluating system performa:p.ce, working
directly with a real user population composed of Lehigh faculty and staff. These methods are now under
development.
A full discussion of the text-processing procedures
now being used in LEADER is contained in earlier
publications. 1 For ease of reference, however, we shall
begin our detailed description of system design and
operation with a brief synopsis of theory and software
development up to the present time.
Retrieval theory
The theoretical basis of LEADER may be described
as a deductive framework for the operationsgoveming
the retrieval of those messages whose content is such as
to make them measurably relevant to negotiable inquiries. By "message" we mean a well-formed as....c;:emblage of lexical items constituting a sentence, a document, a portion of a document, a set of sentences culled
from several different documents, a document description, a bibliographic reference, a citation, and so On.
This liberal interpretation of "message" is intended to
reflect the wide variety of useful retrieval responses to
(a) different types of inquiry, and (b) different stages
within the negotiation of a given inquiry. It is thus
apparent that the LEADER system is designed to incorporate within a single framework the hitherto
separate functions of document retrieval, data retrieval,
reference retrieval, etc. There are numerous reasons for
combining these different functions into one integrated
mechanism, some of which relate to improved hardware
and software, while others arise from changes in design
philosophy. By far the most important reason, however,
is that inquiries put to a retrieval system are not all of
the same kind, so that it is imperative to build into the
system as many different types of response as there are
different kinds of inquiry. Many requests, for example,
require a substantial degree of retrieval completeness,
while others may be met by rather simple lists of document references. A flexible response capability has therefore been programmed for LEADER, permitting the
system to vary its several outputs as the user specifies,
whether during the COlL~e of negotiating a single inquiry, or responding to different types of request.
The theoretical framework of LEADER is described
in a series of pUblications2 dating from 1962. At its most
recent stage of evolution, the theory defines three major
groups of operations leading to a corresponding identification of three main system components, each with its
own sub-theory. We distinguish a text-processing, a
connectivity, and a decision component, and describe
below how their theoretical foundations were established.
Text-processing subtheory
This subtheory formalizes a set of procedures for
assigning content-indicating symbols to documents.
In the LEADER system such symbols are termed
"characteristics", while the activity of assigning characteristics to documents is called "characterization".
Since we characterize a document in order to mention
the topics it deals with, it is clear that the objective of
text-processing is to enumerate a set of topic-denoting
expressions for every document in the collection. By
limiting our scrutiny to noun-phrases, it is possible to
identify topic-denoting expressions as those substantival expressions occumng as arguments of logical relations. For example, the sentence "The four-group constitutes a subgroup of the tetrahedral group" expresses
a binary relation whose arguments are "the fourgroup" and "the tetrahedral group". We say that these
two noun-phrases are topic-denoting expressions, or
that the sentence deals with the topics of the fourgroup and the tetrahedral group.
In the complete text-processing theory of LEADER,
each document is regarded as a complex of ordered
sentences, every one of which must be reduced to its
underlying logical relations. We say that any sentence
expressing a relation has canonical form, or is a canonical component, and call the process of reducing an
From the collection of the Computer History Museum (www.computerhistory.org)
The LEADER Retrieval System
English sentence to ~ts canonical components canonical decomposition. An algorithm has been written to
identify the canonical components of every input
sentence of a document, and to isolate all noun-phrases
occurring as arguments of such components. These
noun-phrases are potential document characteristics
whose actual selection is subject to rules described in
another section.
Connectivity subtheory
The second subtheory is concerned with (1) a class
of operations defined on characteristics to establish
their interconnections, and (2) a class of operations
defined on documents and their characteristics. The
former class of operations serves to formalize the concept of association, here interpreted as the calculable
relatedness of topic-denoting characteristics. The latter
class of operations formalizes the notion of affiliation
between characteristics and documents. This concept is
used to measure the strength of the connection between
a given characteristic and any document to which it has
been assigned.
Because the output of text-processing functions as
the input data for connectivity, it is necessary to provide a link between the respective subtheories. This can
be accomplished by stipulating that terms are connected
at the first level if they occur as the arguments of a
relation.
A criterion was established for measuring the degree
of connection between each term used to characterize
a document and the document itself. All characteristics
are assigned w~htB relative to their documents, represented by the entries of a chalacteristic by document matrix. Multiplication of this matrix by its transpoSe yields a characteristic by characteristic matrix
defining connections between characteristics via one
document.
This matrix is partitioned into submatrices defining
ge'nera. A genus is defined as a connected component
of a graph of characteristics, and is therefore made up
of characteristics every pair of which is connected by
a sequence of edges in the graph.
Each genus submatrix is converted to a transition
matrix, and it is shown that properties of connectivity
within general have an appropriate model in the theory
of ergodic M,arkov chains.
A description of the procedures used to establish connections is provided in a later section.
Decision subtheory
The goal here is to construct a theory of the opera-
449
tions governing the deductive and associative liaisons
between document representations and inquiries. This
particular theory has gone through many stages, and is
still in fact being rather vigorously studied. A major
reason for continued investigation is that our conception of the decision component had been revised several
times in response to the increasing possibilities of ever
more sophisticated retrieval. Thus, starting with the
relatively straightforward problem of constructing a
model for subject document retrieval, we have gone
on to examine associative models, and have finally
arrived at the stage whereby the retrieval process is
defined as an interactive procedure between the user
and the system in which there are many different
levels and types of response, ranging from a simple
enumeration of bibliograppjc information, at the lowest
level, to a full negotiation of a complex query involving
browsing, text-display, and question-answering, at the
highest level. As the decision component has become
more complex and sophisticated, so its corresponding
subtheory has passed through many stages of mathematical development, the most significant of which
have been described in earlier publications. 3
System implementation
Hardware configuration
The LEADER system has been implemented on an
IBM 1800 Process Control computer. The process controller (cpu) has a 16K, 4#L second main core memory
and a 2310 disk controller with two 2315 disk cartridges that provide over a million words of online peripheral storage. An IBM 1442 card reader/
punch, an 1816 typewriter console and a 2260 cathode
ray tube terminal are used for basic system I/O operations.
The IBM 1800 Process Control system operates
under TSX, an IBM-written time-sharing executive
with real time capability. I t is through TSX that
LEADER functions.
Users can communicate with the LEADER System
in either of two ways. The first method is via an 1816
typewriter-like terminal that transmits data at about
15 characters per second over the 1800's standard
data channel. It is an on-site (local) device. The second
method is via a 2260/2848 Video Display-Controller
which transmits data at 240 characters per second over
a selector channel. The 2260 CRT is a video display
terminal with a 960 character display buffer (screen)
and a typewriter-like console for data entry . Up to
eight 2260 CRTs can be attached to the 2848 display
controller at various remote locations.
From the collection of the Computer History Museum (www.computerhistory.org)
450
Spring Joint Computer Conference, 1969
1. General papers on information retrieval, docu-
Software
The LEADER System software can be divided into
three main categories based on producer and function.
These are:
IBNI supplied system software
LEADER System interface software
LEADER System operational software
The IBIVI supplied software consists of the standard
software packages available with the 1800 Process
Control system.
The LEADER System interface software performs
two basic functions. The first is a system interface routine that links TSK and the 2260/2848 display controller. I t services hardware interrupts generated in
the LEADER System's operational environment.
The second is a program interrupt-servicing routine
that supervises all LEADER operations requested ,by
the user.
The LEADER System operational software consists
of three classes of routines designed to perform the
functions required by each of its three components.
They are:
Text Processing Component Program Package
Text Entry Compiler (LETEXT)
String :\1anipulation Compiler (LECOM-II)
Syntactic Analyzer (LEGRA]\'[)
Connectivity Component Program Package
Connectivity Operations Supervisor
Connectivity Matrix Routines
Retrieval File Generation Routines
Decision Component Program Package
Request Analyzer
Request Negotiator
Display Generator
Data base
The Lehigh University Center for the Information
Sciences maintains a literature collection in information
science and engineering as both an experimental data
base and an active reference source for faculty and
students. At present there are approximately 3000
documents in the collection, and our intent is to limit
the size to 10,000 items. However, we wish to make the
collection highly dynamic so that is will continue to be
substantively useful and meaningful to researchers and
students. This will be accomplished by input filtering
and item deletion. At the input stage, only highquality documents dealing with topics of substantial
interest will be accepted for processing and inclusion.
The criteria for selection are as follows.
mentation, and computer appreciation are excluded, unless there is a specific reason for acceptance, such as basic policy statement (e.g.,
Weinberg Report) or a particularly cogent
description of the field, or a general paper in
which specific important data are included.
2. The following areas are included in the collection, with particular emphasis on research,
experimentation and systems analysis.
(i) Automatic indexing and abstracting;
(ii) Syntactic analysis (but not when the
orientation is exclusively mechanical
translation) ;
(iii) Logical and mathematical studies of
retrievai, reievance, indexing, etc. ;
(iv) Basic systems studies, including costs,
major system studies, parallel system
problems, compatibility, library automation;
(v) Behavioral studies of users, questions,
effect of information on management
decision, the research program, and on
engineering processes.
(vi) Programming languages, particularly
symbol manipulation languages;
(vii) Pertinent reviews and bibliographies;
(viii) Peripheral material considered pertinent;
automata studies, mechanical trapslation, self-organizing systems, cognitive
processes, neurophysiology, linguistics;
(ix) Education of personnel in information
- handling and the information sciences.'
With respect to item deletion, user reaction is the
basis for upgrading the quality of the collection. Since
it is hardly meaningful to experiment with documents
that real users find of little interest, we shall delete items
found to be of minimal value, and replace them by documents suggested by users.
These filtering, deJetion and replacement operations
are not only necessary for quality control, but ensure
that the experimental collection is non-static. Collections that are sealed off against new entries misrepresent
real-world conditions, and our approach to system design is to deal exclusively with dynamic collections,
whose membership is subject to continuous modification.
Since the literature collection is, to the best of our
knowledge, the only formal and accessible set of highquality documents devoted exclusively to information
science and engineering, it has certain unique features.
Not only are faculty, research personnel and students
From the collection of the Computer History Museum (www.computerhistory.org)
The LEADER Retrieval System
using the collection for experimentation, but the literature itself is of substantial interest to them. In order to
make the literature available for reference purposes, it
is controlled by conventional library methods. All
incoming documents are classified and shelved, and
author and subject indexes are maintained.
451
N I A of P Euclid's A algorithm U capable A of P
treating V L three A or P more A equations N I A
in P one A process N IV, CC but C nevertheless C
introduces VX many A extraneous A factors NIA
(2)
where the category symbols stand for the following:
Text entry and processing
Each dOCllillent selected for inclusion in LEADER is
entered in its full-text version. The instrument for such
text-entry operations is a data-editing, non-numeric,
and non-interpretive compiler known as LETEXT (see
section on Software).
Text is entered by direct keying from an 1816 terminal or a 2260 ORT and stored on disk files. Input
creates both visual and computer records, and extensive
editing and error correction features are available.
Text entry and editing in LETEXT are background
jobs on the IBM 1800 time-shared computer, so that
the machine-readable data base of LEADER may be
continuously augmented while retrieval operations are
in progress.
The procedure for analyzing the sentential structure
of documents entered into LEADER involves three
distinct steps. The first is a dictionary look-up program
for identifying the key functor words and phrases of any
sentence. The dictionary currently contains approximately 2,000 entries, including such items as suffixes,
prepositions, prepositional phrases, conjunctions, conjunctive phrases, auxiliary verbs, and so. When found,
each functor word is replaced by a specially coded
syntactic category symbol.
The second program assigns syntactic category
symbols to the remaining words of the sentence on the
basis of a context-sensitive analytic grammar.
The third program reduces the strings of categories
produced by the first two programs to their canonical
components. These, as previously explained, are strings
of categories having the form of a sentence expressing a
logical relation.
To illustrate the procedure down to the decomposition stage, consider the following example of ~ input
sentence:
Brown's algorithm is a generalization of Euclid's
algorithm capable of treating three or more equations in one process, but nevertheless introduces
(1)
many extraneous factors.
The output of the dictionary look-up procedure for
this sentence is:
Brown's A algorithm U is V a ART generalization
A
U
V
ART
NI A
P
VL
N IV
CC
C
VX
adjective
unknown
verb
article
noun or adjective
preposition
present participle or gerund of unknown
verb
noun or verb
comma
conjunction
present or past tense of transitive or intransitive verb.
The second program replaces the unknowns of (2)
with calculated category symbols and resolves multiple
category assignments such as N I A. We have:
Brown's A algorithm N is V a ART
generalization N of P Euclid's A algorithm N
capable A of P treating V L three A or P more A
equations N in P one A process N, CC but C
nevertheless C introduces YX many A extraneous
(3)
A factors N
The third program decomposes (3) into the following
canonical components:
(i) Brown's A algorithm N is V a ART generalization N of P Euclid's A algorithm N
(ii) Brown's A algorithm N is V capable A of P
treating V L three A or P more A equations W
in P one A process N
(iii) Brown's A algorithm N introduces VX many A
extraneous A factors N
(4)
This example shows clearly how the phrase "Brown's
algorithm" has three logical occurrences, although it
appears just once in the original sentence.
When every sentence of an input document has been
reduced to its canonical components, another program
selects document characteristics from such components
and assigns numerical weights to selected characteristics. Since all potential document characteristics are
noun-phrases occurring as arguments of the relations
From the collection of the Computer History Museum (www.computerhistory.org)
452
Spring Joint Computer Conference, 1969
expressed by canonical components, they are extremely
easy to identify.
Each selected characteristic occurring as an argument
of an n-termed relation R is said to h.ave n lines of con~
nection to R. A characteristic-t is said to have m lines
to connection to a document D if m is the sum.of all
lines of connection between t and the relations of D in
which t appears as an argument.
The weight of a characteristic relative to a document
is the sum of its lines of connection to the document.
Thus, if a characteristic t occurs once as the argument
of a two-termed relation in a given document D, and
at another time as the argument of a three-termed
relation of D, th~ weight of t relative to D will be five.
In support of this measure, it is essential to realize
that any document is a coherent complex of assertiO!ls.
Its function is not to enumerate haphazard and disconnected pieces of information, but to give an organized
account of its subject matter. That is, a document fits
together concepts embodying knowledge of a field of
inquiry. The connectivity of terms via predicates contributes greatly to a document's coherence. Such connectivity is, in fact, basic to all other types of connectivity.
In summary, the procedure operates on the full
text of documents written in technical English, reduces
each text sentence to a string of syntactic categories,
resolves each category-string into the set of its canonical
substrings, identifies potential document characteristics
within the canonical substrings, and finally assigns a
weight to each characteristic relative to its parent document.
Once a document has been processed by LEADER,
it is analyzed on a sentence-by-sentence basis for characterization and connectivity decisions. That is, the
structure of each sentence is explored to determine
which of its noun phrases qualify as potential document characteristics. The output of this analysis is a
set of weighted source-derived noun phrases that occurred in referential positions within their respective
sentences. The phrases are sorted, merged and then
entered into a term-document affiliation matrix. By
performing appropriate matrix operations on the affiliation matrix these noun phrases may be grouped into
distinct sets of mutually related terms known as
"genera." Further processing of these genera can b3
done to remove low-entropy terms, hence further rzfining the quality of the document characteristics.
Files
The corpus actually used in the LEADER System
consists of approximately 1000 documents, each of
which is analyzed by the fully automatic text processing
and connectivity procedures described above. The
text of every document is reduced to its major assertions, i.e., sentences making assertions about the topics
denoted by the document's characteristics. Such sentences may be grouped by function, if desired. For example, descriptions of results form one group, while
descriptions of experimental procedures form another.
It can easily be shown that the procedures of canonical
decomposition assist in defining such groups. The
ability to identify and retrieve such sentences is, therefore, an important feature that has been incorporated
in the LEADER System.
The analyzed and reduced text of each document is
placed on the IB11 1800's disk storage cartridges along
with the files required for ret:rieval operations on the
domain of documents, their respective analyzed texts,
their characteristics, and the connections between
characteristics. These files contain bibliographic data
on documents, including the storage locations of the
individual sentences of their text, and describe the
nature and strength of the connections among characteristics, components of characteristics, genera of
characteristics, and documents.
To generate the retrieval files, we make use of the
matrices produced by the connectivity procedures described above. Output from the text processing operations is in the form of an alphabetically sorted set of
source derived noun phrases with a record fannat as
shown in Figure 1.
Document
Number
Source
Derived
Phrase
Weight of
Phrase
Relative to
Document
Sentence
Numbers
in which
Phrase occurs
Figure 1-Text processing output record fonnat
These phrases are used to generate the term-document
matrix on the one hand, and to build the Source Derived Phrase Dictionary. Its record format is similar
to that given in Figure 1, except that identical phrases
are merged relative to a given document. See Figure 2.
Weight
of Phrase
Source
Relative
Phrase Derived Document to DocuNumber Phrase Number ment
Sentence
Numbers
in which
Phrase Doc.
Sent.
Occur No. Wt. Nos.
Figure 2-8ource derived phrase dictionary record format
From this file and information contained in the partitioned term-term matrix, we obtain the Noun PhraseWord Profile Dictionary. Each phrase is reduced to its
non-trivial word components. Along with each such
From the collection of the Computer History Museum (www.computerhistory.org)
The LE...A~ERRetrieval System
word, the phrase number of the phrase from which the
word was obtained and the genus number of that
phrase are combined to for a Temporary Word File
with records as shown below.
Phrase Number
Genus Number
Word
453
Citation File
Document File
The Special Topics File is a topic-oriented file made
up of individual sentences from various associated
documents. The remaining files are generated directly
at input of the document via LETEXT.
Figure 3-Temporary word file
Interactive retrieval
These records are then sorted first alphabetically by
word, next by genus number, and finally by phrase
number. Then this sorted list is merged into the Noun
Phrase-:-Word Profile Dictionary whose record format
is shown in Figure 4.
Word
Genus
Number
Phrase
Numbers
Genus
Number
Phrase
Numbers
Figure 4-Noun phrase-word profile dictionary
The Temporary File can then be discarded. Finally,
we generate the Phrase Affiliation File which is derived
directly from the partitioned term-term matrix. It
contains the genus number, phrase number and affiliation value.
The genus number is assigned to each of the submatrices in the partitioned submatrix. Within each
genus (submatrix), the column and row entry of the
submatrix is a source-derived phrase, while the component entry within the submatrix is the term-term
affiliation value derived for each pair of terms within
the given genus. The result of merging this information is the Phrase Affiliation File. See Figure 5.
Affiliated
Affiliated
Genus
Phrase String
Affiliation String
Affiliation
Number Number Number Value
Number Value
Figure 5-Phrase affiliation file
The end product of the sorting and merging of the information contained in the matrices are the three main
retrieval files described above. They are:
Source Derived Phrase Dictionary
Noun Phrase-Word Profile Dictionary
Phrase Affiliation File
There are four other data files used by LEADER.
These are:
Special Topics File
Author/Title Information File
LEAD ER is designed to encourage user interaction
with the structured material of a corpus of scientific
or managerial data so as to maximize the influence of
information flow on decision making. The data entry
procedures are sufficiently general to accommodate
several different types of data base, provided only that
each consist of well-formed English sentences. In 00dltion, the response capability is flexible enough to
p 3rmit retrieval ranging from the enumeration of simple
bibliographic data, on the one hand, to full-text display,
on the other.
The extent to which information flow contributes to
decision-making is certainly affected by the ability of an
information system to adapt itself to a user's needs.
I t is for this reason that retrospective literature searches
are no longer sufficient in information retrieval. It is
now necessary to develop the framework and experimental procedures requisite for a true interaction between user and store. Most experimental work to date
looks upon both the inquiry and the relevance of answers
a 3 single events. We think this is a mistake and that an
inquiry is merely a micro-event in a shifting, adaptive
process. It is not a command, as in conventional search
strategy, but rather a description of an area of doubt
in which the question is open-ended, negotiable and
dynamic. The immediate goal of the LEADER system
is thus to provide a facility that will, within feasible
and practical limits, offer the user a range of experimental configurations which he can amend or add to
as necessary. Its long range function is to design and
test techniques that will allow inquirers to be instructed
in the system, to browse, to query, to be interrogated
by the system, and to be shown various strategies for
search.
To begin a retrieval dialogue, a user enters a preliminary search description (via the 2260 CRT console) in
the form of a set of declarative English sentences.
LEADER's syntactic analyzer then reduces these
sentences to a set of noun phrases which were found
to be in referential positions within the sentences.
N ext, the noun phrases are reduced to non-trivial
component words. This is done because it is rather unlikely that a noun phrase presented by a user will precisely match any of the noun-phI"ase8 in LEADER's
From the collection of the Computer History Museum (www.computerhistory.org)
454
Spring Joint Computer Conference, 1969
source derived Noun Phrase Dictionary. Thus, each
component word of each noun phrase derived from the
request is looked up in the Noun Phrase-Word Profile
Dictionary to determine its respective genus association and phrase affiliation, if any exists. The results
are then merged by genus into maximal sets of ranked
noun phrases, if any exist, affiliated with the request
component words. LEADER's response will be a set of
noun phrases within a single genus that contains the
maximum nunlber of request component words.
Consider the following simple example. Suppose
that a user's request was reduced to the noun phrase
Genus
3
Phrase
finite Nlarkov chains
finite chains
ergodic ~vIarkov chaL.!S
finite Boolean lattice
~'Iarkov Processes
Phrase
Genus
bounded finite space
6
finite set
finite Markov chain
Genus
by LEADER. This phrase would then be further reduced to the three word components
finite
Markov
chain
Next, LEADER would look up each of these word
components in the Noun Phrase-Word Profile Dictionary and the Noun Phrase Dictionary. Suppose the result
was as follows:
Word
Component
Genus
Phrase
finite chain
finite Boolean lattice
{
finite l\1arkov chain
Phrase
No.
1
chain
~ll
/
3
4
5
finite chain
finite :?vlarkov chain
{
ergodic :Markov chain
3
4
chemical chain
chain reaction
{finite molecular chain
8
9
10
1
Phrase
No.
6
7
No. of
words
1
1
Frequency
1
1
Phrase
No.
10
8
No. of
words
2
1
1
Freqrency
1
1
1
9
3
2
2
1
1
to continue with. LEADER will now proceed to obtain
all phrases in genus 3 that are affiliated with the two
selected phrases. This information is obtained from the
Phrase Affiliation File and the Noun Phrase Dictionary.
The two sets of affiliated phrases are then merged and
ranked by affiliation value. The results may be as
follows:
Affiliated
Phrases
Value
finite l\'Iarkov chain
10
unique probability
vector
5
finite state automata
4
vector algebra
3
ergodic :Markov chain ergodic Markov chain
9
unique probability
vector
7
finite state automata
4
context-free grammars
4
Phrase Phrase
No.
3 finite lVlarkov chain
4
Merging the phrases by genus and ranking them by
the number of word components each phrase affiliated
phrase contained, and then by frequency of occurrence
of the affiliated phrases, the result would be as follows:
Frequency
3
finite Markov chain
ergodic Markov processes 4
3
finite ~larkov chain
ergodic :;Vlarkov chain
{
Markov Processes
No. of
words
3
2
2
1
1
In this case, the phrases in genus 3 would be presented
to the user as a response. It would then be up to the
user to select or reject this output based on what it is
that he is looking for. (If it were the case that the user
rejected all of these phrases, LEADER would present
him with the phrases from the next highest ranked
genus, 11.) Let us assume that the user is satisfied with
the phrases in genus 3. He may at this point choose to
see any docll.."1lents associated with these phrases
(or any combination of them), or he may continue his
negotiation to further clarify his request. Suppose he
selects the phrases
2
J?o~nded finite space
lhrnte set
Markov--3
11
Phrase
finite molecular chain
chemical chain
chain reaction
Phrase
No.
3
1
4
2
5
From the collection of the Computer History Museum (www.computerhistory.org)
The LEADER Retrieval Systenl
Science Foundation for support under Grant No.
Gl\-668 of the work on which this paper is based.
The merged and ranked response would be
Affiliated Phrase
A.ffiliation Value Sum
unique probability vector
finite Markov chain
ergodic Markov chain
finite state automata
context free grammars
vector algebra
12
10
9
8
4
'1
u
Again, the user may decide to look at various document
sets associated with any combinations of these phrases
he may choose. Or he may again continue with these
phrases to further define his request. The user also has
the option of returning to any preceding step in his
request negotiation at any time, or he may choose an
entirely new search direction. The important point is
that there is a continuous dialogue between the user
and the LEADER system allowing the user to become
familiar with LEADER'S file organization, and manipUlating it as he wishes.
When the user decides he has reached a point at
which he would like to see some documents, LEADER provides him with a very flexible documental unit
display feature. He may choose to see: author/title
information
citations
full text
In the case of full text display, the user may view the
complete document, if he desires, in a continuous
display, or he may choose to see all those sentences in
a given document which contain one or more of the
phrases he has selected earlier. Finally, he may choose
to see sentences from each document presented to him
containing associated phrases. Thus the user has complete browsability free of language and hardware use
restrictions.
ACKNOWLEDGMENT
Grateful acknowledgment
455
IS
made to the National
REFERENCES
(i) D J HILLMAN
Characterization and connectivity
Document Retrieval Relevance and the Methodology
of Evaluation National Science Foundation
Grant ~o GN-451 Report ~o 1 May 24 1966
(ii) W R HILTON D J HILLMAN
The structure of LECOM
Ibid June 29 1966
(iii) D M REED D J HILLMA~
M icrocategorization for text-processing
Ibid July 7 1966
(iv) D M REED D J HILLMAN
Canonical decomposition
Ibid August 12 1966
2
(i) Problems, systems and methods
Study of Theories and Models of Information Storage
and Retrieval Grant G24070
The National Science Foundation August 3 1962
(ii) The Boolean algebra model
Ibid
(iii) A positive model for systems of special classification
Ibid August 29 1962
(iv) New foundations for retrieval theories
Ibid August 12 1963
(v) Positive models of retrieval systems as species of
logical algebras
Ibid August 23 1963
(vi) Retrieval systems for non-static document collections
Ibid September 26 1963
(vii) Graphs and algorithms for term-relations
Study of Theories and Models of Information Storage
and Retrieval
The National Science Foundation Grant ~o GN-283
July 30 1964
(viii) The structure of document relations
Ibid August 25 1964
(ix) Topology and document retrieval operations
Ibid July 1 1965
(x) The formal basis of relevance judgments
Mathematical Theories of Relevance with Respect
to the Problems of Index:ng National Science
Foundation Grant No GN-177 July 9 1964
(xi) An algorithm for document characterization
Ibid March 12 1965
;~
Cf. Referen~e 2.
From the collection of the Computer History Museum (www.computerhistory.org)
From the collection of the Computer History Museum (www.computerhistory.org)