Skip to main navigation Skip to search Skip to main content

On extracting structured knowledge from unstructured business documents

  • Gaurav Pandey
  • , Rakshit Daga

Research output: Contribution to conferencePaperpeer-review

3 Scopus citations

Abstract

Efficient management of text data is a major concern of business organizations. In this direction, we propose a novel approach to extract structured knowledge from large corpora of unstructured business documents. This knowledge is represented in the form of object instances, which are common ways of organizing the available information about entities, and are modeled here using document templates. The approach itself is based on the observation that a significant fraction of these documents are created using the cut-copy-paste method, and thus, it is important to factor this observation into business document analysis projects. Correspondingly, our approach solves the problem of object instance extraction in two steps, namely similarity search and then extraction of object instances from the selected documents. Early qualitative results on a couple of carefully selected document corpora indicate the effective applicability of the approach for solving an important component of the efficient text management problem.

Original languageEnglish
Pages155-162
Number of pages8
StatePublished - 2007
Externally publishedYes
EventIJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007 - Hyderabad, India
Duration: 8 Jan 20078 Jan 2007

Conference

ConferenceIJCAI 2007 Workshop on Analytics for Noisy Unstructured Text Data, AND 2007
Country/TerritoryIndia
CityHyderabad
Period8/01/078/01/07

Fingerprint

Dive into the research topics of 'On extracting structured knowledge from unstructured business documents'. Together they form a unique fingerprint.

Cite this