Skip to content

What this cannot do

If you read one page before buying, read this one. Nothing here is legally required boilerplate; it is the list of things we know are wrong or unproven with our own product.


It is not a diagnosis, and it is not close to one

It cannot tell you whether you have any condition. It will not name one. It has no clinical validation, no clinician in the loop, and no cohort behind it.

The report describes patterns in text you wrote. Turning a pattern into a meaning about a person is a job for a person — ideally one who is qualified and who knows you.

If something in the report worries you, the useful next step is a conversation with a professional. It is not a second report.


We have run this on one person

N = 1. One real ChatGPT export has been through the full pipeline end to end.

The central question this project was built to answer is: does it tell two people apart, or does it flag everyone the same way? We cannot answer that yet. Answering it requires at least two subjects, and we have one.

That is not a small caveat. A tool that produces a plausible, well-cited, internally consistent report about everyone who uses it would look exactly like this one does today. We do not think that is what we have — but we have not proved it isn't, and we would rather sell you a report with that written on the box.


The numbers we measured, including the bad one

Run the same corpus through twice and you do not get identical output — the models are not deterministic, and chasing determinism would cost more than it buys. So we measure agreement between runs instead.

WhatAgreementRead this as
Which sessions are notable0.78Fairly stable. Two runs mostly point at the same places in your history.
Which pattern type a session gets0.50Weak. Two runs agree about where to look far more than about what to call it.
The synthesis stage0.83Stable, after we found and fixed a structural bug that had it at 0.11.

The 0.50 is the honest problem on this page. It means the label attached to a stretch of your history is roughly a coin-flip between runs, even though the location is reliable. This is why the report gives you the quotes, the everyday explanation, and the counter-argument rather than a confident category — the underlying material is more trustworthy than any name we put on it.


Our extractor has a known blind spot

When we tested the extraction model against hand-built cases, it scored 78% overall — and 0 out of 3 on one specific marker type: how a person responds when the assistant contradicts them.

That matters more than the aggregate suggests, because the report treats the mix of pattern types as meaningful. An extractor that cannot see a type produces a profile where that type is nearly absent, and a reader interprets absence as a fact: "this person doesn't dig in when challenged." The tool's limitation gets rendered as a statement about a human being.

So: any pattern type whose measured recall is below the floor is named in your report, and the corpus-level pass is forbidden from making claims about it. Uniform mediocre recall is safer than lopsided excellent recall, and we choose models accordingly.


A guardrail that leaks about one time in six

The report is instructed never to use clinical or diagnostic register about you. We measured how well that instruction holds by generating the same report repeatedly and having independent judges check the text.

It leaks in roughly 1 of 6 draws. Not zero. An instruction-based guardrail is probabilistic and fails silently — which is why there is now a panel of four independent judges checking the finished report as a structural gate, rather than trusting the instruction alone.


What we would need to see to stop selling this

These were written down before we had results, so that a disappointing outcome could not be quietly re-interpreted as a good one. We publish them because a claim about your own falsifiability is worthless if you keep it private.

We stop, and write up the negative result, if:

  1. It does not discriminate — subjects get effectively identical profiles, and tightening the pipeline does not change that.
  2. Findings cannot be traced — a large share of candidate findings fail the citation gate, meaning the model is inventing them.
  3. Reproducibility collapses — agreement stays below 0.5 after addressing the cause.
  4. It causes harm without being useful — a subject reports the output was distressing and told them nothing, or a report reads as a de-facto diagnosis despite the guardrails.
  5. Consent or privacy fails — anyone withdraws, or the redaction pass is found leaking personal information.

Number 3 came close. Synthesis agreement measured 0.11 at first. The cause turned out to be structural rather than a prompt problem, and the fix took it to 0.83 — but it was a genuine near-miss, and the kill-switch was worded in a way that would have pointed at the wrong remedy.


Things it structurally cannot see

  • Anything you did not type. Voice conversations, other assistants, other apps, and your entire life away from a keyboard are all invisible. A chat log is a keyhole.
  • Why. It sees that a topic returned eleven times. It has no idea whether that is rumination, a research project, or your job.
  • Context that changes everything. A bereavement, a deadline, a new diagnosis, a move — the export contains only what you happened to type near those events.
  • Selection bias in the medium itself. People bring problems to an assistant and solutions to nobody. Chat logs are systematically skewed toward unfinished and unresolved things. The report is reading a biased sample and cannot correct for it.

You were not the only author of your own messages

This is the limitation we are least able to do anything about, and it has become more serious rather than less.

Every message you wrote was a reply. It was shaped by what the assistant had just said — its register, its enthusiasm, how much it agreed with you, whether it asked a follow-up. People mirror their interlocutor, and an assistant tuned to be warm and agreeable is a strong one to mirror. So the transcript is not a record of you thinking. It is a record of you thinking with that particular thing, in that particular period.

Assistant memory tightens this into a loop. Once a system synthesises a profile from your earlier messages and conditions later replies on it, the sequence becomes:

  1. You write the kinds of things people write to an assistant — stuck, curious, at 1am.
  2. It remembers that slice and treats it as who you are.
  3. Its replies are shaped by that slice.
  4. You respond to those replies, writing more of the same kind of thing.

Memory did not fix the sampling problem. It industrialised it. What we read as your pattern may be a groove the two of you wore together.

Three consequences we want you to hold while reading your report:

  • A pattern we report may be an artifact of the loop rather than a fact about you. We cannot separate the two from your side of the transcript alone, and we do not pretend to.
  • Your corpus is not a stable object. Two exports from the same person at different times sit on different memory states and different model versions. Some of what looks like change over time is the assistant changing, not you.
  • This runs deeper than our own exclusion rule. We bar the model that wrote your assistant turns from judging them. But through memory and mirroring, that model also shaped your user turns — so the corpus is more entangled with it than "don't let it grade its own homework" quite covers.

There is a second pipeline in this project that reads the assistant's turns instead of yours, precisely because the human-side screener cannot see this. It produces population-level metrics and no document about any individual. It is built and almost entirely unmeasured, so nothing in your report comes from it — but that is where an honest answer to this would have to come from.


If reading it is hard

The report quotes you back to yourself, and that can land harder than expected — particularly if your history contains a bad period. If your export includes messages about self-harm or ending things, the report will say so before you reach them, because being surprised by your own worst night is worse than being warned.

If any of it is heavy, go through it with someone you trust. If you are in crisis right now, please use one of these — this document is not a substitute for a person.