Skip to content
Iknow

From Overwhelming to Actionable: A Text Mining Proof-of-Concept for Voice-of-Customer Data.

Pharmaceutical company · Pharmaceuticals & Biotechnology

Evaluating Entity Extraction, Sentiment Analysis, and Medical Taxonomy Integration Against Real Internal Reports

Executive summary

Iknow worked with Company J’s Scientific Affairs to evaluate how advanced text-mining software could identify and extract insights from an overwhelming amount of customer information. Scientific Affairs generates, packages, and disseminates medical and clinical information supporting more than 40 commercial brands. Each of Company J’s information-gathering groups — its field-based Scientific Affairs Liaisons, Customer Contact Center, Medical Communications Group, and Medical Education Group — prepared monthly Voice of the Customer (VOC) reports, and Scientific Affairs wanted to know whether text analysis software could streamline the analysis and dissemination of that steadily growing body of information.

Over a focused one-month engagement, Iknow selected the SAP BusinessObjects Text Analysis suite to analyze three months of unstructured VOC reports for a specific therapeutic area and product, processing input files in Microsoft Word, PowerPoint, and PDF formats. Iknow tested the product’s out-of-the-box functionality, then developed Scientific Affairs-specific customized rules through an iterative process, incorporating findings from Iknow’s earlier Knowledge Audit project and integrating MeSH (Medical Subject Headings), the National Library of Medicine’s own controlled medical vocabulary, to build a domain-relevant lexicon. Iknow also created custom rules for sentiment analysis and for flagging sentences describing potential next steps or research opportunities. The proof of concept demonstrated that the software’s entity extraction engine, combined with a customized medical taxonomy, could successfully extract relevant terms, phrases, and sentiment from internal reports with a good balance of results specificity.

Background & context

About the Client

Company J’s Scientific Affairs Division generates, packages, and disseminates medical and clinical information on more than 40 commercial brands. It operates with roughly 350 employees who cover the United States.

Industry Context

This proof of concept came at a pivotal moment for the underlying technology: SAP had acquired Business Objects in early 2008, folding in Business Objects’ own prior acquisition of Inxight, whose SmartDiscovery and ThingFinder technology became the foundation of SAP BusinessObjects Text Analysis. For pharmaceutical Scientific and Medical Affairs functions, Voice of the Customer reports synthesizing feedback across many channels represented a rapidly growing but manually intensive analytical burden. Text mining offered a path to systematically surface patterns and sentiment that would otherwise require exhaustive manual review — work with genuine clinical and regulatory stakes in a domain where emerging safety or efficacy signals matter. Grounding a proof of concept in an authoritative, domain-specific vocabulary like MeSH, rather than relying on generic language processing alone, reflected a broader industry recognition that specialized, safety-critical medical content demands more rigor than off-the-shelf text analytics typically provide.

Current Situation

The objective of this proof of concept was to investigate the effectiveness of text analysis software in streamlining the analysis and dissemination of VOC information. The scope was deliberately focused: analyze three months of reports for a specific therapeutic area and product, and evaluate how well the software could identify and extract key terms and phrases. Iknow selected the SAP BusinessObjects Text Analysis suite to analyze the unstructured text in internal reports, extract critical information, and categorize the findings.

Problem / challenge

  • An overwhelming, growing volume of customer information. Scientific Affairs’ various groups generated an overwhelming amount of Voice of the Customer information that was becoming difficult to analyze and disseminate manually.
  • No proven approach to automated insight extraction. No proven method existed for using automated text mining to extract meaningful insights from Scientific Affairs’ own unstructured internal reports.
  • Generic text mining tools poorly suited to medical content. Scientific Affairs needed a vocabulary and taxonomy that reflected its specialized medical and clinical domain, not a generic, off-the-shelf text-mining approach.
  • No evidence base for a larger technology investment. Scientific Affairs needed real evidence, not assumptions, about whether a text analysis platform would work before committing to broader deployment.

Project objectives

  • Investigate the effectiveness of using text analysis software to streamline the analysis and dissemination of VOC information.
  • Analyze three months of reports for a specific therapeutic area and product to test the software’s real-world performance.
  • Evaluate how well the software could identify and extract key terms and phrases from the reports.
  • Identify where else text mining and text analysis software could be applied within Scientific Affairs.

Iknow’s approach

How Iknow Structured the Work

Iknow structured the engagement as a focused, scoped proof of concept — establishing a baseline with out-of-the-box software functionality, then iteratively customizing rules and lexicons against real Scientific Affairs’ content — reflecting Iknow’s AI proof-of-concept methodology for validating a technology approach with evidence before any production investment.

Key Activities & Decisions

  • Platform selection. Iknow selected the SAP BusinessObjects Text Analysis suite to analyze Scientific Affairs’ unstructured internal reports, extract critical information, and categorize the findings.
  • Multi-format file analysis. Iknow analyzed input files in Microsoft Word, PowerPoint, and PDF formats, reflecting the real mix of document types that Scientific Affairs’ groups actually produced.
  • Iterative custom rule development. Iknow used the product’s out-of-the-box functionality as a baseline, then developed Scientific Affairs-specific customized rules through an iterative process of analyzing the product’s outputs.
  • Integration of prior Knowledge Audit findings. Iknow incorporated the findings and recommendations from its earlier Knowledge Audit project to ground the proof of concept in Scientific Affairs’ actual content landscape.
  • MeSH lexicon integration. Iknow incorporated MeSH (Medical Subject Headings), the National Library of Medicine’s controlled vocabulary thesaurus used for indexing PubMed articles, to create a relevant, domain-specific lexicon.
  • Sentiment and opportunity-detection rules. Iknow created custom rules for sentiment analysis, covering both positive and negative product mentions, and for identifying sentences that discuss a potential next step or an opportunity for further research.

Stakeholders & Collaboration

Iknow served as prime contractor, working directly with Scientific Affairs on this focused, single-month proof of concept, building on the trust established during Iknow’s earlier Knowledge Audit engagement with the organization.

Challenges & how Iknow overcame them

Testing Generic Software Against Highly Specialized Medical Language

Off-the-shelf text analysis software is built for general business language, not the specialized, safety-relevant vocabulary of pharmaceutical Scientific Affairs content. Iknow addressed this by incorporating MeSH, the National Library of Medicine’s own controlled medical vocabulary, directly into the proof of concept, giving the software a domain-appropriate lexicon rather than relying on its general-language capabilities alone.

Balancing Precision Without Missing Real Insights

Rules tuned too loosely would bury genuine insights in noise, while rules tuned too tightly risked missing meaningful terms and sentiment altogether. Iknow addressed this through an iterative rule-development process, repeatedly analyzing the product’s outputs and refining custom extraction and sentiment rules until it achieved, in its own assessment, a good balance in results specificity.

Results & impact

Quantitative Outcomes

  • Scope of analysis: Three months of reports analyzed for a specific therapeutic area and product.
  • Document formats handled: Three input file formats processed: Microsoft Word, PowerPoint, and PDF.
  • Custom rule sets delivered: Custom rules developed for entity extraction, sentiment analysis, and next-step and opportunity identification.
  • Expansion opportunities identified: Six additional application areas identified for broader text mining deployment: taxonomy creation, metadata creation, search and browse improvement, content management tagging, call center data analysis, and VOC report analysis.
  • Engagement duration: One-month assignment.

Qualitative Outcomes

Iknow demonstrated that the SAP BusinessObjects Text Analysis product’s entity extraction engine, combined with a customized medical taxonomy, could successfully extract relevant terms, phrases, and sentiment from Scientific Affairs’ internal reports, and demonstrated how rules and lexicons were created and how stemming could identify additional meaningful terms. The proof of concept gave Scientific Affairs evidence-based confidence that text mining could extend well beyond VOC reports — into taxonomy creation, metadata creation, search and browse improvement, content management tagging, and call center data analysis — and could integrate with business intelligence tools to create online dashboards for executive reporting.

Timeline to Impact

Within the single-month engagement, Iknow delivered a working, evaluated proof of concept spanning custom rules, medical lexicon integration, and sentiment analysis, along with a documented view of where the technology could extend across Scientific Affairs’ broader operations.

Iknow’s capabilities demonstrated

Core Skills

  • AI and text mining proof-of-concept development
  • Automated content classification and entity extraction
  • Taxonomy and medical lexicon integration
  • Sentiment analysis rule design

Methods & Frameworks

  • Iterative rule development and tuning
  • Domain-specific lexicon integration (MeSH)
  • Proof-of-concept evaluation against a defined scope

Technologies & Tools

  • SAP BusinessObjects Text Analysis suite
  • MeSH (Medical Subject Headings); entity extraction and stemming

Put this experience to work on your problem.

Much of our work never reaches the website. Book a call, tell us your sector and we will walk you through the engagements that map to yours.