A Duplicate Chinese Document Image Retrieval System Based on

IDEA GROUP PUBLISHING
14 Chan,701
Chen,
& Ho
E. Chocolate
Avenue, Suite 200, Hershey PA 17033-1240, USA
16*'$!
Tel: 717/533-8845; Fax 717/533-8661; URL-http://www.idea-group.com
Chapter II
A Duplicate Chinese
Document Image Retrieval
System Based on Line
Segment Feature in
Character Image Block
Yung-Kuan Chan, National Huwei Institute of Technology, Taiwan
Tung-Shou Chen, National Taichung Institute of Technology, Taiwan
Yu-An Ho, National Taichung Institute of Technology, Taiwan
ABSTRACT
With the rapid progress of digital image technology, the management of duplicate
document images is also emphasized widely. As a result, this paper suggests a duplicate
Chinese document image retrieval (DCDIR) system, which uses the ratio of the number
of black pixels to that of white pixels on the scanned line segments in a character image
block as the feature of the character image block. Experimental results indicate that
the system can indeed effectively and quickly retrieve the desired duplicate Chinese
document image from a database.
INTRODUCTION
With the widespread popularity of electronic documents as well as electronic
books, more and more people have started reading these new sorts of publications. Most
ThisCopyright
chapter appears
the book,
Systems
and Content-Based
Retrieval
edited by
© 2004,in Idea
GroupMultimedia
Inc. Copying
or distributing
in print or Image
electronic
forms, without
written
Sagarmay
Deb. ofCopyright
© 2004,
Group Inc. Copying or distributing in print or electronic forms
permission
Idea Group
Inc. isIdea
prohibited.
without written permission of Idea Group Inc. is prohibited.
Duplicate Chinese Document Image Retrieval System
15
of the electronic books and electronic documents are stored in an image format. As for
traditional newspapers or magazines, they can also be transformed into digital image
formatted data by using the tool like a scanner, and then saved in a computer for the
purpose of data backup. Meanwhile, by combining with the characteristic of rapid data
processing of a computer, the data of the duplicate documents can be quickly acquired
and retrieved.
Because the technology of Optical Character Recognition (OCR) has been improved
to a degree, a document image can be transformed to become the data with Big5 code
format, and only the Big5 codes are stored so as to reduce the required memory space.
However, using the technique of OCR may cause some errors in document contents
because of recognition failure. Moreover, in order to preserve the original contents and
states of a document, transforming a document into an image and storing the image in
a database is sometimes necessary.
The primary function of a duplicate document image retrieval (DDIR) system is to
transform the data in document format into those represented in digital image format. This
transformation would be completed by using the tool like a scanner. Then, the DDIR
system will store these images and their corresponding features in a database for data
backup purpose. Here the document image is called a duplicate document image. When
intending to retrieve a duplicate document image, users can also use the tool like a
scanner to input the first several text lines of the original document into the system, and
to create a query document image and figure out the feature of the image. The DDIR
system finally transmits the users the duplicate document image whose image feature is
similar to that of the query document image.
In the aspect of extracting the feature from a document image, many techniques have
been proposed (Caprari, 2000; Doermann, Li, & Kia, 1997; Peng et al., 2001). Caprari (2001)
used the method of document template to measure the overlapping ratio of black pixels
and white pixels between two document images when their two templates overlap.
Generally, the higher the overlapping ratio is, the more similar these two documents are.
Doermann, Li and Kia (1997) encoded the heights of English letter shapes and used these
codes as the feature of the English document image. Meanwhile, by comparing and
matching these feature codes, the duplicate document images that are stored in the
database can be searched and retrieved. Besides, Peng et al. (2001) applied the block sizes
and locations to be the feature of a document image, and suggested a component block
list matching technique based on this feature for the duplicate document image retrieval.
The features mentioned above are only suitable for stating the characteristic of an
English document image. The characteristics of Chinese characters are different from
those of English ones, and the strokes and shapes of Chinese characters are much more
complicated than those of English characters. Therefore, this paper suggests a line
segment feature, which is appropriate for stating the characteristics of the Chinese
character image blocks in a Chinese document image, and applies this feature to construct
a duplicate Chinese document image retrieval (DCDIR) system.
The next section will give a brief review of some related works. The following section
will describe the line segment feature in a character image block proposed in this chapter.
Then, the chapter will delineate how to make use of this proposed feature to create a
DCDIR system. This is followed with a statement on the experimental and analyzed
results. This is followed by a discussion on the future trends of the duplicate document
image retrieval system. At last, the conclusions will be given.
Copyright © 2004, Idea Group Inc. Copying or distributing in print or electronic forms without written
permission of Idea Group Inc. is prohibited.
8 more pages are available in the full version of this
document, which may be purchased using the "Add to Cart"
button on the publisher's webpage:
www.igi-global.com/chapter/duplicate-chinese-documentimage-retrieval/27053
Related Content
3D DMB Player and Its Reliable 3D Services in T-DMB Systems
Cheolkon Jung and Licheng Jiao (2012). Depth Map and 3D Imaging Applications:
Algorithms and Technologies (pp. 434-450).
www.irma-international.org/chapter/dmb-player-its-reliable-services/60279/
Spontaneous Facial Expression Analysis and Synthesis for Interactive Facial
Animation
Yongmian Zhang, Jixu Chen, Yan Tong and Qiang Ji (2011). Computer Vision for
Multimedia Applications: Methods and Solutions (pp. 20-37).
www.irma-international.org/chapter/spontaneous-facial-expression-analysissynthesis/48311/
Enhancing Robustness in Speech Recognition using Visual Information
Omar Farooq and Sekharjit Datta (2012). Speech, Image, and Language Processing
for Human Computer Interaction: Multi-Modal Advancements (pp. 149-171).
www.irma-international.org/chapter/enhancing-robustness-speechrecognition-using/65058/
Relevant Feature Subset Selection from Ensemble of Multiple Feature
Extraction Methods for Texture Classification
Bharti Rana, Akanksha Juneja and Ramesh Kumar Agrawal (2015). International
Journal of Computer Vision and Image Processing (pp. 48-65).
www.irma-international.org/article/relevant-feature-subset-selection-fromensemble-of-multiple-feature-extraction-methods-for-textureclassification/151507/
Science of Emoticons: Research Framework and State of the Art in Analysis
of kaomoji-type Emoticons
Michal Ptaszynski, Jacek Maciejewski, Pawel Dybala, Rafal Rzepka, Kenji Araki and
Yoshio Momouchi (2012). Speech, Image, and Language Processing for Human
Computer Interaction: Multi-Modal Advancements (pp. 234-260).
www.irma-international.org/chapter/science-emoticons-research-frameworkstate/65062/