Skip to content Skip to footer
Blog

ISPOR 2026: Beyond the Battle of the Bots, the Future of AI in Evidence Generation

MadeAi | ISPOR 2026: Beyond the Battle of the Bots, the Future of AI in Evidence Generation Angeline Dhas July 21, 2026
MadeAi | ISPOR 2026: Beyond the Battle of the Bots, the Future of AI in Evidence Generation

AI in Evidence Generation was the quiet headline at ISPOR 2026 in Philadelphia this May. Most exhibit-hall chatter, though, was still about which chatbot gives the “smartest” answer. MadeAi presented four posters at ISPOR 2026. They covered literature review methodology, GenAI validation, drug repurposing, and social media listening for CSL Behring. Together, they tell a simple story: raw model intelligence is only step one. What really decides whether a health economics and outcomes research (HEOR) team can trust an AI-generated evidence table is different. It comes down to the workflow, the governance, and the human oversight built around the model.

This post walks through the findings from MadeAi’s ISPOR 2026 research. It looks at why “battle of the bots” thinking misses the point for pharma, biopharma, medtech, and CRO teams, and what a defensible framework for AI in evidence generation actually looks like.

Why “Battle of the Bots” Misses the Point

Ask two general-purpose language models the same clinical question, and you can easily get two different, equally confident answers. That inconsistency is harmless for casual research. It is a different story for a systematic literature review (SLR). An SLR has to survive scrutiny from a Health Technology Assessment (HTA) body, a notified body, or an FDA reviewer, and confident guesswork does not hold up there.

MadeAi‘s research frames the real comparison differently. It is not “Model A versus Model B.” It is unstructured intelligence versus structured, governed intelligence. Left on its own, a general-purpose model tends to treat every study design the same way. A randomized controlled trial, a network meta-analysis, and a case report all get read identically, with no domain-specific workflow behind them. These models often miss data buried in tables, figures, and graphs. Yet these sections frequently contain the most clinically meaningful findings.

By contrast, a specialized, purpose-built platform closes each of these gaps. It adds a dedicated workflow for each study type. It also supports structured capture of tables and figures, full traceability, and enterprise-grade governance. More importantly, the foundation model is treated as one ingredient in a larger system, not the finished product. That distinction matters enormously once a review has to hold up under audit.

CapabilityGeneral-Purpose Model AloneStructured, Governed AI Platform
Study-type handlingReads all designs identicallyDedicated workflow per study type (RCT, NMA, case report)
Multimodal dataFrequently misses tables, figures, and graphsStructured multi-step capture of tables, figures, and graphs
ReproducibilityInconsistent outputs on identical inputsConsistent, reproducible outputs
AuditabilityNo audit trailFull audit trail and traceability
Deployment riskExternal API dependency, limited governanceEnterprise-grade, secure deployment with built-in governance
Regulatory defensibilityNot HTA-defensibleDefensible to HTA bodies, payers, and clinical guidelines

This table sums up the foundation behind every other finding MadeAi presented at ISPOR 2026. It is also the reason AI in Evidence Generation has to be treated as a systems question, not a model-selection question.

A Next-Gen Framework for AI-Augmented Literature Reviews

The centerpiece poster, A Next-Generation Framework for AI-Augmented, Submission-Grade Literature Reviews,” compared five real-world regulatory literature reviews. These supported Clinical Evaluation Reports (CERs) and Periodic Safety Update Reports (PSURs). One review ran as a fully manual baseline. The other four followed an AI-augmented methodology at every stage. That covered protocol development, database searching, deduplication, title and abstract screening, full-text screening, extraction, evidence tables, and PRISMA flowchart generation.

The results stood out for one main reason. Time savings kept climbing even as the workload scaled up dramatically.

ProjectTotal Volume (Articles)AI AccuracyHuman HoursTime Savings
Manual Project120N/A710%
During AI Adoption26588%13017%
Post-AI Adoption (Project 3)13090%6022%
Post-AI Adoption (Project 4)2,30089%61055%
Post-AI Adoption (Project 5)80086%29039%

Article volume grew nearly 19-fold, from 120 to 2,300 articles. However, time savings did not shrink under that load. They actually grew alongside it. That is the practical definition of scalability for AI-powered literature review services. A platform’s efficiency gains should hold up as the evidence base grows. Ideally, they should improve rather than remain limited to small pilot datasets.

Two details matter most for anyone evaluating evidence generation services that pharma teams might adopt. First, AI accuracy stayed consistently in the 84% to 90% range across protocol development, screening, and extraction. That is high enough to meaningfully cut manual effort, but it was never treated as a substitute for human sign-off. Second, human reviewers still adjudicated screening decisions, verified extracted data, and ran final quality control on every single project. The framework did not remove the reviewer. It just removed the drudgery around the reviewer.

The AI-augmented review process, at a glance:

  1. Protocol development. AI drafts inclusion/exclusion criteria and search strategy; a reviewer validates and approves.
  2. Search and deduplication. Automated across databases, with duplicates removed without manual Excel work.
  3. Screening. AI proposes inclusion or exclusion at both title/abstract and full-text stages; humans adjudicate every conflict.
  4. Extraction and synthesis. AI drafts evidence tables; reviewers verify the data behind them.
  5. Reporting. PRISMA flowcharts and final reports generate automatically, with a fully traceable audit trail underneath.

AI-Augmented

AI-Augmented Review Process

Where Human Judgment Still Rules: The Three-Layer Oversight Model

One diagram from MadeAi’s ISPOR 2026 sessions is worth pinning to a wall. Called “Where Do We Draw the Line,” it sorts every task in an evidence synthesis pipeline into one of three layers. The sorting depends on how much interpretive weight each task carries.

LayerOversight ModelExample Tasks
Layer 1: MechanicalAI leads, human verifiesDatabase search, deduplication, title and abstract screening, data extraction, PRISMA compliance
Layer 2: InterpretiveAI assists, human directsFull-text eligibility, risk of bias assessment, heterogeneity analysis, GRADE evidence grading, forest plot review
Layer 3: Judgment (Epistemic)Human only, no exceptionsEvidence interpretation, clinical significance, guideline implications, evidence-gap judgment, practice recommendations

Notice what does not appear in Layer 3. There is nothing about running a search, formatting a table, or deduplicating a reference list. What does appear is the genuinely interpretive work: deciding what a body of evidence actually means for patients and clinicians. That line is drawn on purpose. It is the clearest articulation of “human in the loop” to come out of ISPOR 2026’s AI programming this year.

Three Skills Every Reviewer Needs for AI in Evidence Generation

In addition, the same poster set argued that this layered model creates a new skills requirement for evidence synthesis teams, on top of traditional systematic review expertise.

  1. AI literacy. Knowing when and how AI fails, not just when it succeeds.
  2. Prompt engineering. Instructing the model with the same precision a protocol demands.
  3. Collaborative intelligence. Directing the AI like a junior analyst, rather than simply consuming its output.

Reviewers achieve stronger results when they treat AI as a tool to direct. They do not treat it as a black box to trust. The SA38 data above puts those results at 39 to 55% time savings. This distinction is quickly becoming as central to HEOR training as PRISMA methodology itself. It also connects to why PRISMA reporting still matters in systematic reviews and meta-analyses. A structured reporting standard is worth little if the workflow feeding it is not equally structured.

Validating AI Rigorously: The ELEVATE-GenAI Framework

Efficiency numbers only matter if the AI decisions behind them are actually correct. That is why MadeAi’s second ISPOR 2026 poster applied the ELEVATE-GenAI framework, a structured, multidomain validation approach, to its own platform. The team used a 2,302-record Class II medical device dataset, curated and adjudicated by subject-matter experts.

DomainWhat It MeasuresImpact
Accuracy: InclusionRecall / Precision / F1 for including relevant studies83% / 80.7% / 81.9%
Accuracy: ExclusionRecall / Precision / F1 for excluding irrelevant studies96% / 96% / 95.9%
ComprehensivenessOverall article relevance vs. the human golden dataset93%
FactualityConcordance between AI exclusion rationale and human primary exclusion reasons84%

With these results, two patterns stand out here. Exclusion recall, at 96%, ran higher than inclusion recall at 83%. This reflects a deliberately conservative screening posture. The platform keeps uncertain studies for full-text review rather than excluding them too early. In a regulated setting, that is the right way to fail. Missing a relevant study is a far more serious problem than carrying an extra one forward.

The 84% factuality score goes a step further than a simple accuracy check. It asks whether the AI’s stated reason for excluding a study actually matches the reason a human reviewer would give. That reasoning-level check is what auditability requires, and it suggests the platform’s outputs would hold up under a submission audit.

From Screening to Signal Detection: AI-Assisted Drug Repurposing

The third poster moved from methodology validation to a practical use case: finding drug repurposing signals hidden inside routine literature. Using iron supplementation as a proof of concept, we searched PubMed and retrieved 2,140 records. AI-driven deduplication and screening ran with human oversight at every stage. That narrowed the pool to 853 relevant articles at title and abstract screening, then to 46 studies (13 anemia-focused, 33 non-anemia) at full-text.

AI-Assisted AnalysisAI-Assisted Analysis of Therapeutic Benefits & Reported Risks

AI-assisted extraction then surfaced off-label therapeutic signals across cardiovascular disease, chronic kidney disease, pregnancy, inflammatory bowel disease, cancer, and HIV or orthopedic populations. Benefits ranged from improved exercise capacity and reduced hospital stay to better postoperative recovery. Risks, including gastrointestinal effects and hypersensitivity reactions, were reported just as transparently. The entire review, from protocol through synthesis, took only 84 hours, with an AI screening accuracy of 85%.

That speed matters because drug repurposing signals are secondary and exploratory by nature. A subgroup finding here, an incidental outcome there. These are easy for a human reviewer to skim past inside a 2,140-record haystack. An AI-assisted pass can consistently flag those secondary signals, with a subject-matter expert reviewing every flag. That is a genuinely new capability for hypothesis generation, not just a faster version of an old task.

Listening at Scale: AI-Powered Social Media Insights in HAE

The fourth ISPOR 2026 poster, co-authored with CSL Behring, took the same “structured AI plus human validation” approach well outside literature reviews. This time, the focus was social media listening for hereditary angioedema (HAE). Analysts screened 3,069 posts across X, Facebook, Reddit, Quora, patient forums, and YouTube, with subject-matter experts validating every tagging decision before analysis began.

Compared with a prior six-year analysis from 2018 to 2023, the newer two-year window (2023 to 2025) told a mixed story. Disease burden mentions rose from 61% to 87% of relevant posts, and educational unmet needs climbed from 46% to 50%. At the same time, injection-related burden fell from 12% to 3%, and career-related burden eased from 6% to 4%. In short, therapeutic advances seem to be easing some practical burdens for HAE patients, while diagnostic and educational gaps persist. That kind of nuanced, longitudinal insight is hard to capture through clinical trials alone. It also shows how far an AI platform for Life Sciences can extend when the same governance principles are applied consistently.

What This Means for AI-Powered Literature Review Services in Life Sciences

Put the four posters together, and a consistent picture of a mature Life science solution emerges. AI takes on the mechanical, repetitive load: searching, deduplicating, first-pass screening, and drafting extraction tables. Every interpretive and judgment call still stays with a trained human reviewer. That division of labor made a real difference. It let MadeAi’s teams sustain 84 to 90% AI accuracy, cut human hours by more than half on the largest project, and still produce outputs a regulatory reviewer, HTA body, or payer could audit line by line.

For pharma, biopharma, medtech, and CRO teams evaluating AI-powered literature review services, the ISPOR 2026 data offers a practical checklist rather than a marketing claim. Evaluate whether AI accuracy is reported at each stage, review the available audit trail, and identify which tasks remain exclusively human-led regardless of model confidence. It is equally important to confirm whether performance remains consistent as article volume increases, rather than only on small pilot datasets. Vendors that can answer all four with published, poster-backed evidence are the ones worth trusting. MadeAi did exactly that across SA38, SA52, MSR227, and PCR65, building genuinely defensible evidence generation services pharma organizations can rely on for submission-grade work.

The Road Ahead for Evidence Generation Services in Pharma

The next phase of AI in Evidence Generation will likely be judged on different terms than today’s demos. Consistency against a golden, human-validated dataset will matter more than clever answers. MadeAi’s ISPOR 2026 research is an early, published example of that shift already happening inside real regulatory workflows. CERs, PSURs, HTA dossiers, and social media listening reports are moving from manual to AI-augmented without giving up defensibility.

Ultimately, the important takeaway for HEOR, regulatory, and medical affairs teams is simple. Stop asking which chatbot is smartest. Instead, start asking which platform can prove its intelligence is safe to build a submission on. That proof needs an audit trail, a domain-specific workflow, and a documented human-oversight layer. Structure, not raw horsepower, is what separates a promising demo from a defensible dossier.

Author’s Note: This article was supported by AI-based research and writing, with Claude 4.5 assisting in the creation of text and images.

FAQs

It means combining AI with structured workflows and human oversight to accelerate systematic literature reviews, HTA dossiers, and other regulatory evidence deliverables without compromising traceability or accuracy. Rather than replacing reviewers, AI takes on repetitive, manual tasks, allowing experts to focus on critical analysis, interpretation, and scientific judgment.

MadeAi’s ISPOR 2026 data shows AI accuracy consistently in the 84 to 90% range across protocol development, screening, and extraction, with exclusion-decision precision as high as 96%. That accuracy comes alongside mandatory human validation at every stage, which is what makes the resulting evidence tables and PRISMA flowcharts defensible.

A general-purpose model treats every study design the same way, often misses data in tables and figures, gives inconsistent answers to identical prompts, and leaves no audit trail. A specialized platform adds a domain-specific workflow for each study type, structured multimodal data capture, full traceability, and enterprise governance. That is the difference between an interesting demo and a submission-ready output.

Based on the three-layer oversight model from ISPOR 2026, evidence interpretation, clinical significance, guideline implications, evidence-gap judgment, and practice recommendations should stay human-only, with no exceptions. Mechanical tasks like searching and deduplication can be AI-led with human verification.

Across five real-world regulatory reviews, time savings ranged from 17% during initial adoption to 55% once the workflow matured. Those savings held up even as article volume scaled nearly 19-fold, from 120 to 2,300 articles.

The same structured, human-validated approach extends to drug repurposing signal detection from published literature and to large-scale social media listening studies, such as the CSL Behring hereditary angioedema analysis at ISPOR 2026. That shows the governance model generalizes well beyond a single use case.

Smaller companies need not build AI infrastructure in-house. Strategic partnerships with specialized RWE platforms, disease registries, and academic medical centers provide access to data and analytical capabilities at a fraction of the cost of traditional trials. Academic partnerships often operate on shared-value models. Finally, public funding (NIH grants, EU Horizon Europe) supports RWE studies for rare diseases. The barrier is lower than many assume.