From 75 Gigabytes of Scientific Literature to a Reusable Categorization Model: An Enterprise Search Proof of Concept.
Biotechnology company specializing in human therapeutics · Pharmaceuticals & Biotechnology
Benchmarking Auto-Classification Accuracy Across Three Federal Data Sources for a Global Biotech Company

Executive summary
Company R, one of the world’s leading biotechnology companies, wanted to improve its enterprise search capability, with auto-classification — automatically scanning document content and assigning categories and keywords — as one of its top product requirements. Company R wanted to compare the auto-classification results from the enterprise search products it was considering against the results from one of the leading text analysis products, but lacked the internal expertise to conduct that assessment. Iknow was asked to perform the evaluation, selected for its in-depth technical knowledge and experience with many leading enterprise search and text analysis products.
Iknow chose the SAP BusinessObjects Text Analysis Suite for its superior entity extraction and categorization capabilities, analyzing more than 50,000 scientific and technical documents drawn from the Defense Technical Information Center, the U.S. Department of Energy, and the National Library of Medicine’s PubMed database, totaling more than 75 gigabytes of input data. Using SAP BusinessObjects Text Analysis XI 3.0, its embedded Oracle XE database, and the Categorizer and ThingFinder Workbench tools, Iknow performed 15 separate analyses, including taxonomy creation and auto-classification tasks, classifying content into both the PubMed taxonomy and a proprietary Company R taxonomy with greater than 95 percent accuracy while generating a reusable categorization rule set. Company R made an informed purchase of a new enterprise search product based on Iknow’s analysis, and Iknow further recommended integrating a text analysis product with that platform to build an end-to-end automated content acquisition, tagging, and indexing process.
Background & context
About the Client
Company R is one of the world’s largest independent biotechnology companies. Company R’s research and development organization depends on its scientists being able to reliably find and build on prior research across a large and constantly growing body of scientific and technical literature.
Industry Context
For a research-intensive biotechnology company, enterprise search quality depends heavily on how well scientific and technical content is tagged and categorized, since researchers need to reliably find prior research, competitive intelligence, and technical documentation across a huge and growing corpus. By 2010, auto-classification had become a standard evaluation criterion for enterprise search platforms, but vendor claims of classification accuracy varied widely and were difficult to verify independently. Benchmarking enterprise search products against a proven, independent text analysis engine gave Company R an objective way to validate vendor performance before committing to a purchase. Testing against large, authoritative public scientific and technical data sets like PubMed, DTIC, and DOE gave the assessment real credibility, since these were exactly the kinds of content Company R’s own researchers relied on.
Current Situation
Company R wanted to improve its enterprise search capabilities, with auto-classification as one of its top product requirements, and to compare the auto-classification results from the enterprise search products under consideration with those from one of the leading text analysis products. Company R didn’t have the expertise internally to conduct this assessment, so it asked Iknow to perform the evaluation.
Problem / challenge
- No objective way to validate vendor accuracy claims. Company R needed to evaluate auto-classification functionality as a top requirement for its enterprise search purchase, but lacked an objective way to validate vendor claims.
- No internal expertise for the evaluation. Company R did not have the internal expertise to independently benchmark enterprise search auto-classification performance.
- No large-scale, representative test corpus. Any benchmark needed to be tested against genuinely large, representative scientific and technical content, not a small or artificial sample.
- No reusable methodology beyond a single test. Company R needed the resulting categorization methodology to be reusable, not just a one-time analysis.
Project objectives
- Independently benchmark the auto-classification accuracy of the enterprise search products Company R was considering.
- Test that benchmark against a large, representative corpus of scientific and technical documents.
- Classify content into both an established public taxonomy and a proprietary Company R taxonomy.
- Deliver a reusable categorization methodology and rule set, not a one-time analysis.
Iknow’s approach
How Iknow Structured the Work
Iknow structured the engagement around selecting an independent benchmark platform, assembling a large representative data set, running multiple taxonomy and classification analyses, and validating accuracy and reusability — reflecting Iknow’s AI proof-of-concept methodology for validating a technology approach with real evidence before a purchase decision.
Key Activities & Decisions
- Benchmark platform selection. Iknow chose the SAP BusinessObjects Text Analysis Suite to analyze Company R’s content because of its superior entity extraction and categorization capabilities.
- Large-scale representative data set. Iknow received more than 50,000 scientific and technical documents from DTIC, DOE, and PubMed, each data set including a source-specific taxonomy, full-text documents, and abstracts, with total input data exceeding 75 gigabytes.
- Platform configuration. Iknow performed processing and analysis using SAP BusinessObjects Text Analysis XI 3.0, with the embedded Oracle XE database, Categorizer Workbench, and ThingFinder Workbench tools.
- Taxonomy and entity extraction. Iknow used the Categorizer Workbench’s learn-by-example algorithm and rules-based engine to create and maintain taxonomies, and the ThingFinder Workbench to automatically identify and extract entities from the text.
- Multi-analysis benchmarking. Iknow performed 15 separate analyses on the 50,000-plus document data set, including various taxonomy creation and auto-classification tasks.
- Accuracy and reusability validation. Iknow classified content into the PubMed taxonomy and a proprietary taxonomy with greater than 95 percent accuracy, and confirmed the Categorizer’s learn-by-example algorithm automatically generated a reusable categorization rule set.
Stakeholders & Collaboration
Iknow served as prime contractor, working directly with Company R to conduct the evaluation given the company’s own lack of internal expertise in enterprise search and text analysis benchmarking.
Challenges & how Iknow overcame them
Independently Validating Vendor Auto-Classification Claims
Without a credible, objective benchmark, Company R risked selecting an enterprise search platform based on vendor marketing claims rather than verified performance. Iknow addressed this by selecting the SAP BusinessObjects Text Analysis Suite specifically for its superior entity extraction and categorization capabilities and using it as an independent point of comparison against the products Company R was evaluating.
Producing a Benchmark Robust Enough to Trust at Real Scale
A small or artificial test sample would not have given Company R confidence the results would hold up against its real, large-scale scientific content. Iknow addressed this by assembling more than 50,000 documents exceeding 75 gigabytes from three authoritative sources, then running 15 separate taxonomy and classification analyses to validate the results.
Results & impact
Quantitative Outcomes
- Corpus size: More than 50,000 scientific and technical documents analyzed, drawn from three authoritative sources: DTIC, DOE, and PubMed.
- Data volume: Input data exceeding 75 gigabytes.
- Analyses performed: 15 separate analyses performed, including taxonomy creation and auto-classification tasks.
- Accuracy achieved: Greater than 95 percent classification accuracy achieved against both the PubMed taxonomy and a proprietary taxonomy.
- Engagement duration: Two-month assignment.
Qualitative Outcomes
Company R made an informed purchase of a new enterprise search product based on Iknow’s analysis and outputs, direct evidence that the benchmark produced a genuine purchasing decision rather than just an assessment report. Iknow also recommended that the company purchase a text analysis product and integrate it with the enterprise search tool to create an end-to-end automated content acquisition, tagging, and indexing process. The text analysis software would enhance the enterprise search tool through entity extraction, automated summarization, and autoclassification functionality — extending the value of the proof of concept into a broader technology roadmap recommendation.
Timeline to Impact
Within the two-month engagement, Iknow delivered a validated, reusable auto-classification benchmark exceeding 95 percent accuracy, providing Company R with objective evidence to make an informed enterprise search purchase decision and a roadmap for further enhancing that platform with text-analysis capabilities.
Iknow’s capabilities demonstrated
Core Skills
- AI and text analytics proof-of-concept development
- Automated content classification and entity extraction
- Taxonomy design and development
- Vendor-neutral technology benchmarking
Methods & Frameworks
- Large-scale representative corpus benchmarking
- Learn-by-example (LBE) taxonomy and rule generation
- Multi-source data validation (public and proprietary taxonomies)
Technologies & Tools
- SAP BusinessObjects Text Analysis XI 3.0 (Categorizer Workbench, ThingFinder Workbench)
- Oracle XE embedded database
Put this experience to work on your problem.
Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.

