From Workshop to Production: Building a 1,840-Term Taxonomy for Company Z’s Manufacturing Division.
Global pharmaceutical company focused on prescription medicines and vaccines · Pharmaceuticals & Biotechnology
Custom Taxonomy and Auto-Classification Engine Development for Complex Pharmaceutical Manufacturing Content

Executive summary
Following the taxonomy workshop Iknow delivered to its manufacturing division in late 2018, Company Z asked Iknow to lead the full development of a custom taxonomy and auto-classification engine for its technical content and search tools — enabling its scientists and engineers to find relevant content more quickly across a large, complex content base. Working again as a subcontractor to prime contractor CGI, Iknow built a 1,840-term taxonomy spanning 10 facets and grounded in Company Z's organizational structure, drug nomenclature, manufacturing processes, and direct interviews with department managers.
Iknow then engineered and tuned custom auto-classification rules for every facet, validating results against subject-matter-expert manual tagging — first on a set of detailed test documents, then across a high-level review of roughly 7,000 Company Z documents spanning multiple sites and SharePoint libraries — before training Company Z’s own staff on how to maintain the model. The final taxonomy and classification rules were deployed across all of Company Z's SharePoint content.
Background & context
About the Client
Company Z’s manufacturing division is responsible for formulating, packaging, and distributing Company Z’s products to more than 140 markets worldwide, and for coordinating an interdependent global manufacturing network spanning internal sites and contract manufacturing organizations.
Industry Context
This engagement is the natural next step in a common pattern in knowledge management consulting: an initial advisory engagement establishes trust and technical credibility, then leads directly into a larger build engagement once the client is confident in the approach. Months earlier, Iknow had delivered a taxonomy workshop to Company Z’s KM team, critiquing their existing roadmap and demonstrating Iknow’s depth in both taxonomy methodology and pharmaceutical manufacturing content. Based on that demonstrated expertise, Company Z asked Iknow to lead the full build — moving from strategic guidance to a production-ready taxonomy and classification system serving Company Z’s scientists and engineers directly.
Current Situation
Company Z’s internal knowledge management team sought to enhance and apply a custom taxonomy and auto-classification engine to its content and search tools, enabling Company Z’s scientists and engineers to find relevant content more quickly. Company Z’s manufacturing division asked Iknow to lead this effort, citing Iknow’s experience developing complex scientific and engineering taxonomies and its deep domain expertise in pharmaceutical and biologics manufacturing.
Problem / challenge
- Enormous, multidimensional content to classify. Company Z’s content spanned drug programs, manufacturing sites, contract manufacturing organizations, document types, dosage forms, materials, analytical methods, and known manufacturing issues and deviations — dimensions that do not map cleanly to a single, simple hierarchy.
- No existing structured vocabulary for hundreds of manufacturing-specific concepts. Analytical methods and common manufacturing issues and deviations existed as tribal knowledge across departments rather than as a documented, structured vocabulary that could support a search or classification system.
- Classification accuracy had to be demonstrable, not assumed. Given the technical and regulatory stakes in pharmaceutical manufacturing content, Company Z needed measurable evidence that automated classification matched what subject matter experts would tag by hand — not just a taxonomy that looked reasonable on paper.
- The system needed to scale and remain maintainable. The classification model had to work reliably across Company Z’s entire SharePoint content base, spanning multiple sites and libraries, while Company Z’s staff needed to maintain and adjust the model themselves rather than relying indefinitely on outside consultants.
Project objectives
- Build a custom taxonomy model that reflects Company Z’s organizational, scientific, and process structures.
- Ground the taxonomy in real-world knowledge management use cases and the terminology users actually search for and browse with.
- Develop and tune automated classification rules that reliably match content to the correct taxonomy terms.
- Validate classification accuracy against manual tagging by subject-matter experts.
- Train Company Z’s staff to independently maintain and adjust the taxonomy and rules and to prepare the model for full deployment across all Company Z SharePoint content.
Iknow’s approach
How Iknow Structured the Work
Iknow followed an iterative, evidence-based process moving from historical classification review, through use-case interviews and text mining, to taxonomy and rules development, client training, and large-scale validation before deployment.
Key Activities & Decisions
- Review of Historical Classification Schemas. Iknow reviewed Company Z’s functional organization, drug nomenclature and categorization, geographic locations, and contract manufacturing organizations to populate the taxonomy’s core facets. The team studied key process steps in pharmaceutical manufacturing and commercialization, along with the resulting document types, and combined Company Z’s terminology with industry-standard term sets to populate the dosage form and material type facets. The team also built two new facets from internal sources, listing more than 200 analytical methods and 400 common manufacturing issues and deviations.
- Use-Case Interviews. Iknow conducted 15 in-depth interviews with Company Z’s managers across core departments to understand KM use cases and the terminology for search and content browsing. These interviews helped define the appropriate level of granularity for each facet and the term combinations users would actually search for.
- Text Mining. Iknow used Smartlogic’s Text Miner tools on a representative sample of Company Z’s content to surface additional candidate terms, synonyms, and acronyms for concepts already identified through the interviews and document review.
- Taxonomy Model Development and Refinement. Iknow built a first-draft model in the Smartlogic Semaphore Ontology Editor, followed by multiple rounds of editing and review with subject matter experts. The final model contained 1,840 terms organized into a four-level hierarchy across 10 facets, with synonyms, acronyms, and alternative labels for drugs at different development stages, plus more than 200 associative related-term links supporting both classification and later concept browsing for users.
- Classification Rules Development and Refinement. Iknow built and extensively customized XML-based classification rule templates that weighted different kinds of classification evidence — title and body words and phrases, existing metadata, synonym and acronym usage, and negative evidence — per facet, testing results against sample documents throughout.
- Client Training. Iknow prepared and delivered a two-hour interactive online training session covering taxonomy preparation and rules editing and configuration in Semaphore, enabling Company Z’s professionals to make their own future edits using Semaphore’s Task-based individual development environments.
- Testing and Validation. Iknow developed custom auto-classification rules for each facet and iteratively tested them against manual SME tagging on roughly 80 sample documents, using F-score correspondence as the accuracy benchmark. Validation was then extended to a high-level review of approximately 7,000 Company Z documents drawn from multiple sites and SharePoint libraries, followed by final revisions to improve accuracy.
Stakeholders & Collaboration
Iknow worked as a subcontractor to prime contractor CGI, collaborating with Company Z’s internal KM team, department managers across Company Z’s core functions, and subject matter experts who reviewed the taxonomy model and validated classification accuracy throughout the engagement.
Challenges & how Iknow overcame them
Structuring a Genuinely Multi-Dimensional Classification Problem
Company Z’s content needed to be organized along organizational, geographic, process, drug, material, method, and issue-and-deviation dimensions simultaneously — a structure that could easily collapse into an unusable single hierarchy if handled carelessly. Iknow addressed this by organizing the taxonomy into 10 distinct facets, each grounded in a different, complementary source: internal organizational data, process and document-type analysis, a blend of internal and industry-standard terminology, and original research that built two entirely new facets from internal sources.
Proving Classification Reliability at Production Scale
A taxonomy that looked reasonable on paper was not sufficient; Company Z needed proof that automated classification was reliable enough to trust across a large, technically sensitive content base. Iknow addressed this by using F-score correspondence between auto-classification and SME manual tagging as an explicit, measurable accuracy benchmark. They iterated on roughly 80 detailed test documents until results were satisfactory, then scaled validation to a review of about 7,000 documents across multiple SharePoint sites and libraries, identifying and correcting accuracy issues before full deployment.
Results & impact
Operational Outcomes
- Delivered a 1,840-term taxonomy organized into a four-level hierarchy with 10 facets, including two entirely new facets built from internal sources that cover more than 200 analytical methods and 400 common manufacturing issues and deviations.
- Conducted 15 in-depth use-case interviews with Company Z department managers and added more than 200 related-term links across the taxonomy hierarchy.
- Validated classification accuracy against manual SME tagging on roughly 80 sample documents using F-score correspondence, then extended validation to a high-level review of approximately 7,000 documents across multiple sites and SharePoint libraries.
Strategic and Organizational Outcomes
The final taxonomy model and classification rules are being deployed across all of Company Z’s SharePoint content, supporting a planned update to the user interface. Iknow’s two-hour training session equipped Company Z professionals to make their own future edits and adjustments to the model and rules, giving the Company lasting, independent ownership of the system rather than ongoing reliance on outside consultants.
Timeline to Impact
Iknow completed the nine-month build engagement directly following the diagnostic workshop delivered the previous year — together representing a complete workshop-to-production taxonomy lifecycle.
Iknow’s capabilities demonstrated
Core Skills
- Taxonomy development & modeling
- Auto-classification engineering
- Content classification validation & QA
- Client enablement & training
Methods & Frameworks
- Facet-based taxonomy design
- Use-case-driven terminology development
- Text mining for candidate term discovery
- F-score-based classification accuracy validation
Technologies & Tools
- Smartlogic Semaphore (Ontology Editor, Text Miner, classification rules engine)
- XML-based classification rules templates
- Microsoft SharePoint (content deployment target)
Put this experience to work on your problem.
Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.
