Ask a CI lead what their function costs and the answer is usually dominated by subscriptions. That is not automatically wrong — assembling pipeline and trial data independently is genuinely expensive and genuinely undifferentiated. But it is worth being precise about what the money buys, because the market is easy to over-purchase and the overlap between products is larger than the sales process suggests.
What is actually on the market
| Category | Representative products | What it is genuinely good for |
|---|---|---|
| Pipeline and trial | Citeline (Pharmaprojects, Trialtrove), Clarivate Cortellis, GlobalData | Knowing that an asset or trial exists, with structured fields and change history. The base layer most other work sits on. |
| Commercial and forecast | Evaluate, GlobalData, analyst consensus sources | Consensus revenue expectations and comparability across assets. Useful as a reference point, not as a forecast of your specific situation. |
| Prescription and claims | IQVIA, Komodo Health, Symphony and similar | Observed treatment behaviour — what is actually being prescribed, to whom, in what sequence. The only reliable counter to anecdote. |
| Market access | MMIT, policy and formulary trackers | Coverage policy, utilisation management criteria and formulary position by plan. Frequently the missing half of a competitive picture. |
| Regulatory and label | Cortellis regulatory modules, agency sources directly | Approval history, label text and precedent. Note that the primary agency sources are free and often better. |
| Monitoring and alerting | Purpose-built CI platforms, general media monitoring, increasingly LLM-based tools | Catching disclosure you would otherwise miss. Quality varies enormously and the failure mode is volume rather than absence. |
Two caveats on that table. First, product names and modules change frequently. Second, and more importantly for budgeting: this market has consolidated heavily. Several of the names above sit under shared corporate ownership after a decade of acquisitions, which means two subscriptions you think of as independent may be drawing on related underlying data. That is worth checking before treating agreement between two sources as corroboration.
What databases are reliably good at
The honest version: a good pipeline database is an excellent record of what has been disclosed. It will reliably tell you that a trial exists, who sponsors it, what phase it is in, what it says it is measuring, and — the underrated feature — what those fields used to say. Structured change history is the single most useful thing most teams are already paying for and not using.
They are also good at breadth. A human analyst will not find the Phase I programme at a Chinese biotech that nobody has written about in English. A database, imperfectly, often will.
The four things no subscription answers
- Which of these actually competes with us. A query returns forty assets in the tumour type; perhaps six occupy the same line, biomarker subset and combination setting as yours. Getting from forty to six is judgement applied to domain knowledge, and it is the request that was actually made.
- What the result means. A hazard ratio in a field is not an interpretation. Whether a trial's control arm was an acceptable standard of care, whether the subgroup driving the effect is the commercial population, whether the follow-up is mature enough to believe — none of that is a field.
- What the competitor intends. Databases record disclosure. Intent is inferred from patterns in disclosure — an amended enrolment criterion, a quietly dropped cohort, a site expansion in a new region — and inference is not a lookup.
- What was said rather than published. The discussant's framing, the response to a question from the floor, the investigator's view of the toxicity in practice. This is the primary-research half of the problem, and it is covered separately in primary versus secondary intelligence.
The buy / build line
Buy the data layer. It is expensive to assemble, undifferentiated once assembled, and nobody has ever won on having a marginally better copy of the trial registry.
Build the judgement layer. Competitive-set definitions, interpretation, primary access and the maintained view of what matters. This is the part that is specific to your portfolio and the part that decays if unstaffed.
The common failure is the reverse — a comprehensive subscription stack and nobody with the domain knowledge to turn a query result into an answer.
Running an overlap audit before renewal
Worth doing once a year, and it takes about a day:
- List every subscription, its cost, and the named internal owner. Subscriptions without an owner are usually the ones nobody uses.
- For each, write the specific recurring question it answers. If two subscriptions answer the same question, one of them is a renewal candidate.
- Pull actual usage statistics from the vendor. Ask for them by seat. Vendors will provide this and the results are frequently uncomfortable.
- Check what is available free. Trial registries, agency approval and label databases, the Orange and Purple Books, and company filings cover a surprising amount of what is being paid for — see LOE and biosimilar intelligence for how much of that work runs on public sources.
- Identify the questions nothing you hold can answer. That gap is usually the better place for the marginal budget.
On the AI layer
Every vendor in this market now has a natural-language layer over its data, and there is a parallel set of tools offering to monitor and summarise autonomously. The useful ones are genuinely useful: summarising a large volume of disclosure, drafting a first pass, catching things a keyword alert would miss.
The specific risk in a CI context is that these tools produce fluent output regardless of whether the underlying retrieval was right, and CI output is used to make decisions that are expensive to reverse. A summary that quietly omits the competitor programme nobody indexed reads exactly like one that did not. The practical position is to use them for coverage and drafting, and to keep a human accountable for the claim — particularly for anything that will be presented as a competitive judgement rather than a record of disclosure.
A reasonable default
For most teams: one pipeline and trial subscription rather than two, chosen on the quality of its change history as much as its coverage. Claims or prescription data if you have marketed products and a real question about treatment behaviour. Market access data if access is a live competitive variable, which it usually is. Agency and registry sources used directly rather than through a wrapper. And the balance of the budget on the people and primary access that turn all of it into an answer, because that is the part the subscription cannot supply.