Abstract
Efficient management of text data is a major concern of business organizations. In this direction, we propose a novel approach to extract structured knowledge from large corpora of unstructured business documents. This knowledge is represented in the form of object instances, which are common ways of organizing the available information about entities, and are modeled here using document templates. The approach itself is based on the observation that a significant fraction of these documents are created using the cut-copy-paste method, and thus, it is important to factor this observation into business document analysis projects. Correspondingly, our approach solves the problem of object instance extraction in two steps, namely similarity search and then extraction of object instances from the selected documents. Early qualitative results on a couple of carefully selected document corpora indicate the effective applicability of the approach for solving an important component of the efficient text management problem.
| Original language | English |
|---|---|
| Pages | 155-162 |
| Number of pages | 8 |
| State | Published - 2007 |
| Externally published | Yes |
| Event | IJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 - Hyderabad, India Duration: 8 Jan 2007 → 8 Jan 2007 |
Conference
| Conference | IJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 |
|---|---|
| Country/Territory | India |
| City | Hyderabad |
| Period | 8/01/07 → 8/01/07 |
Fingerprint
Dive into the research topics of 'On extracting structured knowledge from unstructured business documents'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver