ChatGPT for homeopathy — what it can and can't do
"ChatGPT for homeopathy" is one of the most-searched practitioner queries about AI in this field. The reason is workable: a general-purpose large language model is free at the entry tier, accepts plain-language case notes, and returns what looks like a rubric list, a remedy suggestion, or a differential within seconds. The narrower question is not whether ChatGPT can produce a plausible-looking homeopathic output, but whether the output is trustworthy enough to inform the second prescription, and whether the case-work cycle — repertorize, materia-medica check, differential, prescribe — survives the substitution. The short answer: ChatGPT is useful for adjacent linguistic tasks and unsafe as the primary repertorization engine.
What ChatGPT is, in homeopathic terms
A general-purpose large language model is not a repertory. It is not a materia-medica database. It does not hold a structured index of rubrics, remedies, gradings, or cross-references the way a repertory engine does. When a practitioner pastes case notes and asks for "the closest rubric in Kent," the model generates a plausible string by next-token prediction; it does not look up the rubric in Kent's Repertory. Two consequences follow. First, the rubric returned may not exist in the cited repertory at all. Second, gradings (italic, bold, plain; first, second, third grade) are routinely invented or omitted because they are not encoded in the model's training distribution at the per-rubric level. Practitioners who lean on ChatGPT as a search surface over the repertory get an answer that reads correct, sounds correct, and is — measured against the printed page — frequently wrong.
This is not a tuning bug. It is a category mismatch between a probabilistic language generator and a discretely indexed bibliographic artefact like Kent's or Boenninghausen's repertory.
ChatGPT versus a repertory engine
| Dimension | ChatGPT (free) | ChatGPT Plus | Similia (Free + Pro) |
|---|---|---|---|
| Repertory coverage | None | None | Kent, Synthesis-class repertories, and premium editions (Murphy MetaRepertory, Complete Repertory, Saine Repertory) |
| Materia medica index | Implicit, ungraded | Implicit, ungraded | Classical libraries plus premium editions (Murphy, Complete, Saine) |
| Source gradings preserved | No | No | Yes, faithful to the cited edition |
| Case storage and timeline | None | None | Unlimited Pro cases, AI-extracted Case Timeline |
| Live consultation capture | None | None | Live Audio Mode (Beta) with consent controls |
| Citation to source rubric | Hallucinated or absent | Hallucinated or absent | Inline to the cited repertory and edition |
| Cost at entry tier (USD) | Free | $20/mo | Free tier; Pro Base $19.99/mo |
| AI provenance and BAAs | OpenAI general model | OpenAI general model | OpenAI for notes/photos, Deepgram for live audio; zero-retention BAAs |
| Suitability as primary repertorization tool | No | No | Yes, with practitioner judgement |
Where ChatGPT helps the workflow
ChatGPT is useful where the task is linguistic rather than bibliographic. Four adjacent uses hold up.
First, translation of patient phrasing into the symptom vocabulary you repertorize from: a patient who says "my head feels like it is in a vice" can be rephrased into the repertory-friendly "head, constriction, as if in a vice" without the model needing to look up Kent.
Second, rough materia-medica summarisation when you already know the remedy — Hahnemann's Materia Medica Pura on Aconitum napellus, condensed into one paragraph, is a tolerable starting point that you then verify against the printed source.
Third, differential drafting at the conceptual level: a query that asks the model to compare the mental picture of Pulsatilla and Phosphorus returns a serviceable outline rooted in Kent's Lectures, which you refine against the page.
Fourth, administrative work — meeting summaries, patient-letter drafts, supplement-interaction quick-checks — where the underlying corpus is broad enough that the model's distribution matches the real world.
In each of these uses the practitioner remains the rubric-mapper and the prescriber. ChatGPT plays a linguistic role, not a clinical one, and the failure modes below do not bite because the model was never allowed to become the bibliographic source.
Where ChatGPT fails — four documented modes
Four failure modes recur and are reproducible on any current build of a general-purpose chat model.
Rubric hallucination. The model returns a rubric phrase that does not exist in the cited repertory, most often inside the "mind" chapter where the rubric tree is densest and the training data thinnest.
Grading drift. Where a rubric does exist, the model reports it as bold when it is plain, or as a three-remedy curiosity when it is in fact a polychrest rubric. Gradings are a Boenninghausen-grade feature of the printed repertory that does not survive an unstructured text corpus.
Differential collapse. ChatGPT anchors on the most-cited polychrests — Sulphur, Phosphorus, Calcarea carbonica, Lycopodium, Pulsatilla — regardless of the case's true centre, because the training distribution over-represents the polychrest corpus.
Confident-tone failure. When the model is wrong, it is wrong in a register that reads as authoritative, with the same fluency it has when it is right. A practitioner who treats the output as a peer-graded suggestion ends up second-prescribing on a hallucinated differential.
Of the four, grading drift is the most clinically expensive: it propagates silently through repertorization and is invisible to anyone who does not cross-check against the printed source. Whether retrieval-augmented chat models, given a verified rubric index, would close this gap remains to be demonstrated; the current generation of consumer ChatGPT does not ship that retrieval surface.
Situating AI tools inside the tradition
A two-century practitioner literature — from Hahnemann's Organon (sixth edition) through Kent's Repertory to the modern sensation-method work of Rajan Sankaran — has built a rigorous case-management methodology whose core unit is the individualised prescription. AI tooling is only as useful as its fidelity to that methodology. Repertorization demands a verified source index, not a probabilistic string generator.
The practical implication: keep the AI on linguistic and administrative tasks; keep the rubric work inside a tool that indexes a real repertory edition, preserves its gradings, and cites the source on every suggestion — for example by searching the repertory by symptom and auditing each suggested rubric against the cited edition before it enters the differential.
References
Hahnemann, S. (1842) Organon of Medicine, sixth edition (posthumous), and Materia Medica Pura; English translation by William Boericke, B. Jain reprint.
Kent, J. T. (1897) Repertory of the Homoeopathic Materia Medica, sixth American edition; Lectures on Homoeopathic Materia Medica, B. Jain reprint.
von Boenninghausen, C. (1846) Therapeutic Pocket Book, translated editions, B. Jain reprint; on rubric grading conventions in the classical repertory tradition.
Sankaran, R. (2005) The Sensation in Homoeopathy, Homoeopathic Medical Publishers, Mumbai.
Verdict
Ready to act on this?