We went into a recent round of user interviews with a hypothesis: researcher distrust AI because it hallucinates, and they want better tools. We expected to hear complaints about cost, reliability, requests for smarter search, maybe some enthusiasm about time savings.

What we actually heard was more clarifying and, in some ways, more sobering.

Our team conducted interviews with researchers across a specialized scientific discipline, deliberately spanning career stage from second-year PhD students to senior faculty with active editorial responsibilities. The goal was to understand how researchers actually find, read, and work with the literature, and to test appetite for various AI-powered research tool concepts.

Levels of familiarity, trust, and enthusiasm varied widely, but five themes surfaced across all our conversations. Those themes are critical perspectives for anyone building AI tools for research.

1. The trust deficit is real, earned, and getting worse

The majority of interviewees raised fabricated citations as a first-order problem without being asked. They volunteered these concerns, often within the first few minutes.

One senior researcher described having an AI tool confidently refer her to a paper that clearly did not exist — not a real paper from the wrong source, but a fully invented reference, plausible-sounding, by a real author working in the right field. Another described being told that his own supervisor had published a 2020 paper. That paper does not exist.

What made the finding sharper was a mechanism one participant described: as models improve, fabrications become more convincing, not less. And when researchers try to verify a suspect citation by searching for it, search engines' AI summaries are increasingly confirming the hallucinations. "AI is reinforcing AI hallucinations," she said. The very verification layer that researchers fall back on is itself being contaminated. Hallucinated papers have been cited in published papers, which reinforces the quality signal to Google Scholar, creating a counterfeit reinforcement loop.

This is a mission-critical integrity problem that is compounding. It came through as the dominant concern across the cohort, even from the two researchers whose distrust took a different form (one found AI too general to be useful; another found it unreliable for tasks that should be trivial). The implication for anyone building AI tools for research is that trust isn't optional—for researchers, it is the core product.

2. The real problem is what lives inside the paper, not finding the paper

Almost every researcher we spoke with uses the same discovery workflow: title, abstract, maybe figures, conclusion. Google Scholar is the nearly universal starting point. It works. Nobody is particularly lost at the paper level. Where they get stuck is inside the paper.

Three researchers described the same workflow tax, independently and without prompting: the specific piece they need — a performance metric, a material property, a figure, a value from a methods section — is buried in prose or supplementary files. To find it, they have to open every candidate paper by hand, scan through, and either find it or discard it. One described this as one of the most time-consuming parts of his work. Another built his own tooling to extract data at scale rather than continue doing it manually. A third described routinely checking supplementary materials on a hunch, having learned that the relevant number was often only there.

Search tools point toward papers. Finding the thing inside them is still manual, page by page. This is the gap that a tool grounded in licensed full text research is positioned to close. It’s a gap that general-purpose AI tools can’t reach, because they don’t have access to the full text and supplementary materials in the first place.

3. They want to ask a question, not craft a query — but the tools that offer that interface can't back it up

One of the clearest patterns across the interviews was the appeal of conversational AI as an interface, even among researchers who distrust it for anything consequential.

As one PhD student put it: AI helps direct him toward the right literature even when he does not yet have a clear mind about the problem or the right vocabulary to search for it. Not knowing the field's precise terminology is a real barrier for early-career researchers, and conversational AI forgives that in a way that keyword search does not.

But the pattern that followed was equally consistent. Generally, AI is useful at the entrance to an unfamiliar field and unreliable as soon as the question gets specific. It handles general orientation. It collapses when a researcher needs a specific absorption spectrum, an evaporation temperature, a precise figure from the literature. One senior researcher described it as "a compass in the dark": a starting point, never the endpoint. Others arrived at the same boundary independently, and two of them hypothesized the reason: models are trained on abstracts, not full text, so they simply do not have access to the specific layer where the information lives.

The researchers want to ask a real question and receive a grounded answer. The tools that offer that conversational interface cannot currently back it up with the specific, verifiable information they need. The interface is right. The architecture underneath it is not.

4. They are judging AI on its weakest form

This is the finding that most reframes the others. With one exception — a power user who builds his own agentic pipelines and sits at the far end of the adoption curve — every researcher we spoke with has formed their view of AI from free consumer tools, standard institutional licenses, or general-purpose chat applications. None of them have used a tool that grounds its answers in licensed, publisher-supplied full text and verifies its citations against a controlled corpus.

Their distrust is earned, but a meaningful part of what they are distrusting is a tier-of-tool problem, not an inherent ceiling on what AI can do. The dominant failures they described (fabricated citations, collapse on specific questions, answers that are general where the researcher needed precise) are architectural problems of tools that perform recall over abstract-level training data with no access to the full-text layer and no citation verification.

The bar that a grounded, citation-verified, full-text tool would need to clear is the bar set by free ChatGPT and institutional Copilot. That is a concrete and reachable bar. For publishers thinking about what to build: the contrast to draw isn’t against other sophisticated tools, it’s against what these researchers are actually using today.

5. No researcher will pay personally; the institution is the only viable path to them

This finding was universal, and it came through without variation across every career stage.

Researchers will not pay meaningful personal sums for specialized research tools. One said he would "find another way" rather than pay. Another noted that students in his country can afford one or two AI subscriptions at most. A third said he would consider one general-purpose tool personally, but any specialized tool has to come through the institution or he does not adopt it. The one outlier in the group — a researcher who does pay for premium tools — proves the rule: he does so because he uses AI at a scale and depth none of the others do, and he is explicit that this is not typical.

The path to researchers runs through institutional licensing. This is not a new lesson for publishers, but it is worth reinforcing as AI tools proliferate and individual subscriptions become the default business model for consumer AI: that model does not reach the research community in any meaningful way.

What this means

We share these findings because we think the publishing community learns faster when we learn together. These are directional signals from a deliberate sample, not representative statistics, but the convergence across independent voices, at different career stages and institutions, on several of these points is enough to ground future explorations and product prototyping.

The researchers we spoke with are using the tools available to them. What they described wanting is something grounded, specific, and honest about the limits of what it knows. That is a reasonable ask, and one the publishing community is uniquely positioned to answer.

1993 1999 2000s 2010 2017 calendar facebook instagram landscape linkedin news pen stats trophy twitter zapnito