health literacy · 2026 · 11 min read
Using AI for health questions: a GP's guide to asking it properly
key points
- Patients are already asking ChatGPT, Claude, Gemini and Copilot about their health. Pretending otherwise helps nobody.
- They work by predicting the next word from text scraped off the internet, and Reddit is repeatedly measured as the single most-cited source behind their answers.
- They are built to be agreeable, so a frightened question tends to get a frightening answer and a hopeful one gets reassurance.
- They are genuinely useful for understanding something you already have, and genuinely dangerous for working out something new.
- The best single test is to ask for a link and then open it. What cannot be linked has not been checked.
- I have written a prompt you can paste in first that forces sources, uncertainty, Australian information and no diagnosis.
More and more, someone sits down and tells me what the AI said before they tell me what is wrong. I am not going to pretend that is a problem I can talk people out of. It is free, it is instant, it is patient with questions, and it never makes you feel silly for asking. Those are real advantages, and general practice should be honest about them.
It is also not a handful of people. A University of Sydney survey of 2,034 Australian adults, published in the Medical Journal of Australia, found that about one in ten had used ChatGPT for health information in the previous six months, and that six in ten of those had asked it at least one question that really needed clinical advice. Use was highest among younger adults, people born in non-English-speaking countries, and people with limited health literacy. That last group is the one I think about most.
What worries me is not that people ask. It is how completely they believe the answer. So rather than tell you not to use it, I would rather teach you to use it the way I would want a colleague to use it: as a decent starting point and a terrible finishing point.
Why does it sound so convincing?
It helps to know what is actually happening under the bonnet. These tools are large language models, or LLMs. Stripped right back, an LLM predicts the next word, then the next, then the next, based on the enormous amount of text it was trained on. It is an extraordinarily good pattern machine. What it is not is a machine that knows things, checks things, or understands that you are a real person with a real body.
So sounding right is the actual job. Producing text that reads like a good answer is what these systems are optimised for, and that is not the same as being correct. Nothing in a calm, fluent, well-organised paragraph tells you which of the two you have been handed.
Then there is where the material comes from. The training text is largely scraped from the public internet, which is books and journalism and reference works, but also an enormous volume of forum posts, comment threads and opinion. It is not curated the way a medical reference is. And when these tools search the web and show you their sources, Reddit turns up more than anything else: one analysis of more than 100 million citations found ChatGPT citing Reddit in close to 60 per cent of its answers in early August 2025, and around 10 per cent six weeks later. A separate study of 30 million sources in March 2026 found Reddit the single most-cited website across ChatGPT, Gemini, Perplexity and Google's AI answers.
Both of those looked at all questions, not health questions specifically, so do not read them as "40 per cent of medical advice comes from Reddit". Read them as this: the sources sitting behind a confident answer are often ordinary people on a forum, the mix changes drastically from one month to the next, and nobody tells you when it does. You cannot see any of that from the answer itself.
On top of all that, they lean towards agreeing with you. Ask "should I be worried about this lump" and you will get a different answer from "this lump is probably nothing, right?" Same lump, same person, different framing, and the answer follows your lead. In the room I can hear when someone is talking themselves into or out of something. A chatbot cannot, and it will happily walk with you in whichever direction you set off.
This has a name in the research, sycophancy, and it has been measured. A Harvard and MIT team asked several models to write a persuasive letter containing a claim that was plainly illogical, and three of the ChatGPT models complied every single time, fifty times out of fifty. My own college puts the everyday version of it well. Launching the RACGP's guidance on these tools, Dr Rob Hosking warned that they "can provide responses that appear authoritative but on review are vague, misleading, or even incorrect".
Even how you type changes what you get back. An MIT team altered patient messages in ways that had nothing to do with the medicine, adding typos, slang, uncertain phrasing, extra spaces, while keeping the clinical facts identical. Every model became roughly 7 to 9 per cent more likely to say the person should just manage it at home. The models also made about 7 per cent more errors for female patients, and were readier to tell women to sort it out themselves. Nothing about the illness changed. Only the typing did.
What is it actually good at?
Quite a lot, and I would rather you used it for these things than not:
- Translating. You have a letter or a result full of words nobody explained. Asking what a term means is a good use of it.
- Understanding a diagnosis you already have. Not working out what you have got. Understanding what you have been told you have got. Those are very different jobs.
- Building your question list. Ask it what a well-prepared patient would ask their GP about this. Turning up with five sharp questions is worth more to me than a printout.
- Rehearsing something awkward. Plenty of people know exactly what they want to raise and cannot find the first sentence. Practising it is a legitimate use.
- Making sense of instructions. Pre-op fasting, a bowel prep, what a form is asking for.
What should I never hand it?
Four things, and I mean these ones plainly:
- A new symptom, to work out what it is. What is missing here is context, and context is most of the job. Your GP knows you, has your history in front of them, can put hands on you and examine you, and has seen a great many people who looked a bit like you and turned out to be fine, and a few who did not. That is clinical acumen. It takes years to build, it is the hardest part of what we do, and it is exactly the part a chatbot has none of. It has your typed paragraph and nothing else.
- Whether you need a hospital. If that question is live, the answer is not in an app.
- Starting, stopping or changing a medicine. Including the supplement you read about, and including stopping something because it listed a side effect.
- A result, read on its own. Again, context is king. Not every abnormal result is bad news: plenty of them are expected, or explained by something we already know about you, or simply how your body runs.
How do I catch it making things up?
Ask for a link, then open the link, then check it says what the chatbot said it says. That is the whole technique, and it catches most of it.
It happens more than you would think. Australian researchers, writing in JMIR Mental Health last November, had a current model write six literature reviews and then checked all 176 references it produced. One in five had been invented outright. Of the real ones, nearly half had errors in them. And the detail that matters most to you: where a made-up reference came with one of those official-looking DOI links, roughly two thirds of those links did open, but landed on a genuine paper about something else entirely. So the link working is not the test. The link saying what the chatbot claimed it says is the test.
The same study found the fabrication got worse on less common conditions: about 6 per cent on depression, nearly 30 per cent on rarer diagnoses. Which is the wrong way round, because the person with the rare diagnosis is exactly the one who cannot easily find good information elsewhere.
If you take only one thing from this article, though, take these two questions. They cost you nothing, they work on every chatbot, and they do more than any clever prompt.
ask it these two things, every time
- "What would change your answer?" A real answer to that names the specific things that would flip it: your age, a medicine you take, how long this has been going on, what a test showed. That list is worth more than the original answer, because it tells you what actually matters and what to go and find out.
- "What have you assumed about me?" It does not know your age, your history, your other conditions or what you take, so it has quietly guessed at all of it. Made to say those guesses out loud, a very certain answer usually deflates on the spot, and you can see straight away how much of it was never about you.
Both questions do the same job: they turn a confident statement back into what it always was, which is a general answer to a general question. That is a perfectly useful thing to have. It is just not the same as advice about you.
Why does it keep giving me American advice?
This one is my own observation from the consulting room rather than something I can point you at a study for, so take it as that. What I keep seeing is medicine names we do not use here, screening ages that are not ours, thresholds borrowed from another country's guidelines, and advice that quietly assumes a health system with no Medicare, no PBS and no GP referral. It makes sense: most of the text these things learnt from was written for an American reader.
What should I never type into it?
Assume everything you type is stored and may be read. Please keep out your full name, your date of birth, your Medicare or health record numbers, your address, and photographs of documents. Do not paste a whole referral letter or a full result: if you want a word explained, paste the word.
A prompt to make it safer
Before anything else: this is not advice, and my actual advice is not to use AI for clinical information at all. It is not a doctor, it has not examined you, and it has no idea who you are.
But I also know how this goes. Many of my patients are already using it, and telling them not to has never once stopped anyone. So rather than pretend otherwise, I decided the next best thing was to put together a set of instructions you can paste in before your question, so that when you do use it, it behaves as carefully as I can make it.
It makes the chatbot use Australian information, ask you questions instead of guessing, give you links you can open, say how confident it is and how old its information is, refuse to diagnose you, keep away from doses, stop agreeing with you automatically, and finish with a list of questions to bring to your GP.
Get the prompt → There is nothing to sign up to, and it is short enough to save permanently in ChatGPT so you never have to paste it again.
None of that makes a chatbot safe. It makes it less wrong, and it makes it much easier for you to see when it is guessing. That is a worthwhile difference, but it is the only claim I am making for it.
Do not take my word for it, run the experiment
You do not have to believe me about any of this, and I would honestly rather you did not. Test it yourself, because it takes about three minutes and it is far more convincing than anything I can write here.
the three minute test
Pick a health question you already know the answer to. Something you have been through, or something we have talked about in the room. Then:
- Open a new chat and ask it plainly, the way you normally would. Read the answer.
- Open another new chat, paste the prompt in first, ask the identical question, and read that answer.
Put them side by side. What most people notice is that the first answer sounded better and the second one was more honest: it asked them something, it admitted which parts it was not sure of, it named what it had assumed, and some of the confident detail from the first answer quietly disappeared when it had to produce a link.
Then, for the real test, go back to the first chat and push back. Tell it firmly that you think it is wrong, even if it was not, and watch how quickly it folds and agrees with you. That is the part I most want you to see for yourself, because once you have watched it happen, you will never quite read one of these answers the same way again.
When to close the app
Some things do not belong in a chat window at all. Chest pain or pressure, sudden weakness or drooping of the face, trouble speaking, difficulty breathing, a severe or sudden worst-ever headache, heavy bleeding, a seizure, sudden confusion, or thoughts of harming yourself. Put the phone down and ring 000, or go to your nearest emergency department. If you are unsure and want to talk to a registered nurse at any hour, healthdirect is 1800 022 222, and here in Victoria NURSE-ON-CALL is 1300 60 60 24. If things feel unsafe, Lifeline is 13 11 14.
And if you have been reading for a while and are no clearer than when you started, book in. Bring what the AI told you if you like, I genuinely do not mind. I would far rather see the printout and talk it through than have you deciding alone.
Where the figures above come from
- How many Australians are doing this. Ayre J, Cvejic E, McCaffery KJ. Use of ChatGPT to obtain health information in Australia, 2024 · Medical Journal of Australia, February 2025. Nationally representative survey of 2,034 Australian adults.
- Agreeing with you (sycophancy). Chen S, Gao M, Sasse K, et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behaviour · npj Digital Medicine, October 2025.
- Typing and phrasing changing the answer. Gourabathina A, Gerych W, Pan E, Ghassemi M, MIT · summary from MIT, presented at the ACM FAccT conference, June 2025. A computer science paper rather than a medical one.
- Invented references. Linardon J, Jarman HK, McClure Z, et al. Influence of topic familiarity and prompt specificity on citation fabrication · JMIR Mental Health, November 2025. Tested on academic writing rather than patient answers.
- My own college's position. RACGP releases new AI guidance · newsGP, July 2025.
- Where AI answers get their sources. Semrush, The most-cited domains in AI · 230,000 prompts and over 100 million citations, July to October 2025. And Search Engine Land, AI search engines cite Reddit, YouTube and LinkedIn most · March 2026. Both are industry analyses, not peer-reviewed research.
All free to read in full, all accessed 20 August 2026. Some tested model versions that have since been replaced.