7/16/2026 · 7 min read
Don't tweet to your critics. Do this instead.
OpenEvidence lost to Nature Medicine. Then it lost the plot on X.

Paul Wicks, PhD
Founder & CEO
→ Proof Points
Quick moment of reflection before we get into it.
We've now released 10 issues! 🎉
Since we have grown a fair bit, I wanted to point newer readers back to three early issues that still come up in conversation.
1) Digital Health is a Gamble, but Where's the Evidence? – This came out of a vote I ran in a room full of mental health founders. What they recognised was that 40 million people a day are using AI for mental health conversations. Their instinct was to be inside those tools as validated modules, not to fight them head-on. I also flagged the "Betrayal" Super Bowl ad as evidence these tools are already being used as therapy whether or not they admit it. I ended up sending that ad to the MHRA (the UK FDA) as an example of exactly that.
2) Your ChatGPT Just Diagnosed You (but it's a fake disease)— this was a fun one. The real finding was that large language models may be trusting the scientific literature too quickly — hoovering up early, unreplicated preprints and repeating them back to users as if they were settled. In normal science, we wait: publish, replicate, review, then rely on it. The machines are skipping to the last step. Hold that thought, because this issue's deep dive is the same problem wearing a different suit.
3. I've helped design clinical trials for 15 years. Then I participated in one— This one was personal. People still reach out about that issue; someone messaged just last week asking how to get involved, and I pointed them to the SmellTaste charity. That one reminded me the point of all this isn't to critique science from the sidelines. It's to be in it.
Right. Let's get to work.
— Paul
Deep diveOpenEvidence forgot the evidence

There are, broadly, two kinds of AI tool in medicine.
The first is general-purpose. The same model that will give you a chicken recipe, read your kid a bedtime story in the voice of Yoda, or advise on solar-panel placement will also, if you ask, answer a medical question.
The second is the specialist: a deep-domain tool trained on a curated corpus of research papers and textbooks, wrapped in guardrails, sometimes wired into an electronic health record, and often restricted to clinicians.
OpenEvidence is the poster child for that second camp.
I first came across the paper that started all this in a departure lounge, killing time on my way to HLTH EU. And something about it tickled my brain; a detail I'll come back to…
Last month, a team at NYU Langone published a Brief Communication in Nature Medicine titled, plainly enough, “General-purpose large language models outperform specialized clinical AI tools on medical benchmarks.” They put OpenEvidence and Wolters Kluwer's UpToDate Expert AI up against three frontier models — GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 — across three benchmarks. The non-specialist frontier models won all three. And on the benchmark built from real, de-identified clinician queries, the specialist tools performed no better than Google's free AI Overview…. yes, the little box at the top of a search page.
Sit with that for a second. If a free search widget matches the tool you're paying a subscription for, someone in hospital procurement is going to ask why the invoice exists (hopefully). But the original study, as first submitted, didn't include that doctor-in-the-loop step at all.
It was the peer reviewers who insisted on it, on the entirely reasonable grounds that letting AI grade AI's homework isn't good enough. If you read the peer-review reports and the back-and-forth between authors, reviewers and editors (and they're public) the whole story is there. One of those reviewers is even named in the paper: Stephen Gilbert, someone I've published with from our Ada days. Small world.
The paper has been read over ~155,000 times and covered in Forbes. In this field, that's enormous.
Now, the response.
OpenEvidence did not do well out of this paper, and had a menu of respectable options: a Letter to the Editor, its own peer-reviewed rebuttal, or simply opening its API so independent researchers could benchmark it properly. Instead it fired off a combative thread on X: what the analyst Sergei Polevikov neatly christened a "twittorial." Some of the technical points were fair; benchmark contamination is a genuine issue, and the authors concede as much in the paper. But the thread also went after the researchers personally: alleging they'd asked for free access to build a better product, and flagging one author's connection to Google—a connection, it's worth noting, that was openly disclosed in the paper's competing-interests section.
Sergei Polevikov, who writes AI Health Uncut, put the asymmetry more bluntly than I could:
So let me do the useful thing and explain why a Letter to the Editor would have been the better move.
First, it becomes part of the permanent record. A letter is bolted onto the original article in the journal and in PubMed, so every future reader sees that a formal challenge exists. A tweet thread evaporates. What if the great debates of our age had been hosted on MySpace?
Second, it's a disciplined format. Roughly 500 words to rebut the main points and to put your own evidence on the table.
Third, it keeps the argument inside the norms of the field, rather than dragging a serious scientific dispute onto a platform that, whatever else you think of it, is not built to be the town square for science.
There's a trade-off: write a letter and the authors get the right of reply, which usually means they get the last word. That ‘s a sure fire path to l'esprit d'escalier, but it’s how we duel in science.
The deeper irony is right there in the name. A company called OpenEvidence met a peer-reviewed critique with no peer-reviewed publications, no disclosed methodology, and a closed API… while lecturing everyone else on rigour.
The Proof Point
When someone calls your evidence ugly, the instinct is to fire back fast. Resist it. The right response to a peer-reviewed paper is not a louder tweet: it's more evidence on the record. Bring data, or bring a letter. Preferably both.
For a picture of how it should be done, here's a small story with a familiar face in it. Five years ago, a German group tested AI symptom-checkers (including Ada) on medical students diagnosing rheumatic diseases. The problem was they mischaracterised Ada as a diagnostic decision-support system, which it isn't, and used a consumer app well outside its intended purpose. Stephen Gilbert and I wrote a Letter to the Editor setting the record straight. The authors got their reply. All of it sits on the permanent record, and the letter has been read around 1,500 times. That's how you correct science. You don't write angry tweets about it. But see if you can spot the passive-aggressive nerd-tropes in zingers like “We read with interest…”, “we thank the authors for drawing this to our attention”, and “for the sake of clarity…”
Reference points:
— Paul
From our deskBeta Testing Claude Science
We are playing with Claude quite a lot at the moment, like everyone, but with a specific angle.
We've been using Claude Cowork to build what are essentially evidence "skills": tools that help a company plan its publications, map a conference strategy, or research a competitor's public claims and assemble a claims bank in a fraction of the usual time. Our current clients get early access to these as we refine them. We're also beta-testers for Claude Science, which plugs into a large number of research databases.
I'll share the useful bits with you subscribers first!
In the meantime, reply and tell me how you're using Claude for research or science. I'm genuinely curious!
Upcoming eventsWhat We're Attending
In early July, I hosted a panel at Health Tech Integrates on partnerships in digital health.
A few themes stuck. The first was that partnerships rarely stay partnerships. Sometimes the company you partner with is the company that eventually acquires you, so it pays to think about where the relationship might end up before you sign. The second was IP: several people made the point that founders can move too slowly out of fear of sharing IP with a potential partner, and that a bit of boldness is usually rewarded. And the funders in the room from Innovate UK described a real shift in the landscape — backing fewer companies with bigger cheques, and pushing them toward revenue and profitability rather than a thousand flowers blooming on public money.
That same evening was the launch of the London chapter of the Hemingway Report, a community of digital mental health operators with branches in London, Melbourne, New York and San Francisco. I dragged along a couple of newcomers I'd met earlier that day to crash the party. We were mildly distracted by England playing DR Congo, but between goals we got into youth mental health, UK reimbursement, and the eternal question of how much evidence is enough. And no, I will not be taking questions on the England v Argentina match.

I'll also be in Boston on 13 September to see a couple of clients I'm proud to work with. It's a tight trip, but if our paths could cross, reply and say hello.

Thanks for reading,
Paul Wicks, PhD
Founder & CEO, ProofStack Health
Move Fast. Prove Things.
P.S. Here are 3 ways I can help you:
Take the Evidence Scorecard Quiz. Answer 15 questions and we’ll send you a personalised report with feedback tailored to your specific needs.
Follow or connect with me on LinkedIn. I publish top resources and in-depth insights related to building your evidence stack.
Book a strategy session. Uncover the gaps in your evidence and marketing in your Digital Health/MedTech startup.
