Extracting Facts from Technical Documents to Expand the Taxonomy.
Global pharmaceutical company focused on prescription medicines and vaccines · Pharmaceuticals & Biotechnology
“Broader and Deeper” Taxonomy Enhancement Through Fact Extraction to Improve Search and Content Browsing

Executive summary
After successful taxonomy development and deployment at Company Z’s manufacturing division, Company Z’s Knowledge Management Center of Excellence wanted to go beyond classification by extracting specific, structured facts directly from documents — including summary content, authors and contributors, and vendors and equipment used in manufacturing — to help scientists and engineers find and evaluate relevant content even faster. Company Z asked Iknow to implement Progress Semaphore’s Fact Extraction module, test it at scale, deploy it to production alongside the existing classification engine, and train Company Z’s KM staff to extend it.
Iknow identified the document types best suited for fact extraction — investigation reports, technical communications, and protocols. Iknow then built and tested Fact definitions in Semaphore against more than 100 documents, successfully extracting summary sentences, author and contributor information, and details such as corrective actions performed. Iknow also used the Facts model to link Company Z’s taxonomy to large external vocabularies of vendors, suppliers, and equipment without adding maintenance overhead to the main taxonomy. Iknow developed search interface mock-ups showing how Facts and extended taxonomy terms could improve result relevance. By the engagement’s close, Company Z was rolling out the new capability across the manufacturing division.
Background & context
About the Client
Company Z is a global pharmaceutical manufacturer. Company Z’s manufacturing division oversees the formulation, packaging, and distribution of its global product portfolio. Company Z’s Knowledge Management Center of Excellence has partnered with Iknow on multiple engagements to build and continuously expand the enterprise taxonomy and search capabilities that serve Company Z’s scientists and engineers.
Industry Context
Classification alone — tagging a document with topic terms — only gets a searcher partway to what they need; it signals that a document is broadly relevant but not what it specifically says. Fact extraction goes a level deeper, using natural language processing to extract specific pieces of information from a document’s text — a summary sentence, the author, the equipment or vendor mentioned, and the corrective action taken — and store them as structured, searchable metadata. Combined with the ability to reference large external vocabularies, such as extensive vendor or equipment catalogs, without folding every term into a core enterprise taxonomy, fact extraction lets an organization meaningfully enrich search without the ongoing burden of managing an ever-growing taxonomy.
Current Situation
During earlier engagements, Iknow built Company Z’s core enterprise taxonomy — more than 1,840 terms across a four-level hierarchy and 10 facets — and deployed autoclassification across all of Company Z’s SharePoint repositories, refining the model through testing, feedback, and analysis of content such as manufacturing deviation reports. Company Z’s KM Center of Excellence wanted to build on that foundation with four objectives: use Semaphore’s Fact Extraction module to extract key facts and extend the taxonomy, test the resulting Facts model on a large document sample, implement it in production alongside the existing classification engine, and train Company Z KM staff to use the fact extraction tool.
Problem / challenge
- Classification alone wasn’t enough. Even with a mature, well-tagged taxonomy, users still had to open and read documents to find specific facts — who wrote it, what it concluded, and which equipment or vendor was involved — that weren’t captured as searchable metadata.
- Some important vocabularies were too large for the core taxonomy. Vendors, suppliers, and equipment model numbers formed enormous, ever-changing lists that would have created significant maintenance overhead if added directly to Company Z’s main taxonomy.
- No proven fact-extraction capability yet existed. Company Z had no model for automatically identifying and extracting facts such as summaries, authorship, or corrective actions from its documents.
- Facts needed a clear path to production and end users. Extracting facts in isolation wouldn’t help anyone unless the capability was implemented in production and reflected in an improved search experience.
Project objectives
- Use Semaphore’s Fact Extraction module to select and extract key facts from Company Z’s documents and to further extend the taxonomy.
- Test the Facts model on a large document sample to ensure accuracy and value.
- Implement the Facts model in production alongside the existing classification engine.
- Train Company Z KM staff in fact extraction for future use.
Iknow’s approach
How Iknow Structured the Work
Iknow structured the engagement into five phases: fact extraction planning and piloting, development and testing, taxonomy extension through external vocabularies, planning for integration with search, and training.
Key Activities & Decisions
- Fact extraction planning and piloting. Iknow reviewed examples of documents believed suitable for value-added fact extraction, especially investigation reports, other technical communications and reports, and protocols, identifying key potential facts including summary sentences and paragraphs, author and contributor information, and vendor and equipment information used in manufacturing.
- Fact extraction development and testing. Using Progress Semaphore, Iknow developed Fact definitions and semantic indicators to process each document, then tested the completed Fact model on more than 100 documents from across Company Z’s repositories. After fine-tuning, the model successfully identified key summary sentences, document authors and contributors, and other helpful text snippets, such as corrective actions performed. The design is extensible to other content types in the future.
- Taxonomy extension. Iknow used the Facts model to link to external taxonomies too large to include in the main taxonomy model, enabling concepts from those external taxonomies to be tagged alongside terms from the main model without added maintenance overhead — significantly expanding tagging for specific vendors, suppliers, equipment items, and model numbers.
- Planning for integration with search. Iknow provided several mock-ups for search interface refinements leveraging the Facts and extended taxonomy, letting users quickly review long results lists and focus on the most useful content, and enabling search based on Facts and extended taxonomy terms with improved relevance. Iknow held several online sessions with Company Z’s IT development team to plan implementation details, including use of the Semaphore API.
- Training. Iknow prepared a comprehensive training session and documentation for fact extraction and taxonomy extension, including how to work with Semaphore tools to develop new Facts and how to test the accuracy of fact extraction.
Stakeholders & Collaboration
Iknow served as prime contractor, working with Company Z’s Knowledge Management Center of Excellence and the IT development team on search integration planning.
Challenges & how Iknow overcame them
Determining Which Facts Were Actually Worth Extracting
Building fact-extraction rules for incorrect information would have wasted effort and cluttered search results without adding real value. Iknow addressed this by running a dedicated planning and piloting phase, reviewing real document examples — investigation reports, technical communications, protocols — to identify concrete, valuable fact types before building any Fact definitions, rather than guessing what might be useful.
Extending the Taxonomy’s Reach Without Expanding Its Maintenance Burden
Vendor, supplier, and equipment vocabularies were far too large and fast-changing to be responsibly folded into Company Z’s core taxonomy. Iknow addressed this by using the Facts model to link to those large external vocabularies instead, keeping the enterprise taxonomy lean while still enabling rich, specific tagging of vendors and equipment.
Results & impact
Quantitative Outcomes
- Facts model testing: 100+ documents tested across Company Z’s repositories.
- Fact types successfully extracted: summary sentences, document authors and contributors, and other helpful text snippets such as corrective actions performed.
- Taxonomy extension: significant expansion of tagging for vendors, suppliers, equipment items, and model numbers via external taxonomy linking.
Qualitative Outcomes
Based on the testing results, Company Z began rolling out the fact extraction capability across the Division. The new capability lets users quickly review long result lists and zero in on the most useful content, and it supports direct searches on Facts and extended taxonomy terms to improve relevance. The Facts model was designed to be readily extensible to new fact types and external taxonomies, and Company Z’s KM staff received comprehensive training and documentation to extend it independently. The work was completed on schedule, and Company Z planned further use and extension of this approach, building on the Division’s multi-year investment in its Iknow-built taxonomy and search infrastructure.
Timeline to Impact
Iknow delivered all five phases within the five-month contract period, running in parallel with Iknow’s study-map taxonomy project for Company Z.
Iknow’s capabilities demonstrated
Core Skills
- Fact extraction & semantic enrichment
- Taxonomy extension via external vocabularies
- Search experience design
- KM training & enablement
Methods & Frameworks
- Document-type-driven fact discovery
- Semaphore Fact definition and semantic indicator development
- Iterative testing and fine-tuning
- Search UI mock-up design
Technologies & Tools
- Smartlogic Semaphore (Fact Extraction module, Classification engine, API)
- SharePoint enterprise search
- External vendor and equipment taxonomies
Put this experience to work on your problem.
Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.
