FR EN
How to test an LLM chatbot or a RAG app in 2026
How to test an LLM chatbot or a RAG app in 2026 | AutomationDataCamp
October 6, 2026 ADC Team 8 min read

How to test an LLM chatbot or a RAG app in 2026: a tester’s checklist

More and more teams ship a chatbot or a “chat with your documents” feature, and testers are asked to sign it off. The usual tools still apply, but the method changes: the same question can get two different answers, the answer depends on documents the model retrieved, and the user can try to talk the model into misbehaving. Here is how we approach it, with the public references behind each check.

Key takeaways
  • Don’t assert exact text: assert properties (required facts, refusal, format, sources) and run each case several times.
  • Test retrieval and generation separately: a wrong answer often comes from the wrong documents, not from the model.
  • Use the OWASP Top 10 for LLM Applications (2025) as a security checklist: prompt injection, sensitive information disclosure, system prompt leakage, excessive agency…
  • Check the AI Act transparency rule: since 2 August 2026, users must be told they are interacting with an AI system.
  • Tools exist: promptfoo (open-source evals and red teaming), Ragas (RAG metrics), and Playwright for the UI.

Why do classic assertions break on a chatbot?

Because the output is not deterministic. Ask the same question twice and you may get two correct answers worded differently, or one correct and one wrong. An assertion like expect(answer).toBe("Your order ships in 3 days") fails on a good answer and tells you nothing about a bad one.

What works better is to assert properties of the answer:

  • it contains the required facts (the delay, the price, the right product name);
  • it does not contain forbidden content (another customer’s data, an internal URL, a competitor’s price);
  • it refuses what is out of scope, instead of inventing;
  • it respects the expected format (JSON schema, length, language);
  • it cites its sources when the product promises it.

Then run each case several times and track a pass rate, not a single green or red. A case that passes 7 times out of 10 is a finding, not a pass.

How do you test a RAG application?

A RAG (retrieval-augmented generation) app first retrieves documents, then generates an answer from them. Test the two steps separately, otherwise you cannot tell where a wrong answer comes from.

  • Retrieval: for a set of questions with known answers, check that the right passages come back. Ragas calls this context recall: how many of the relevant documents were successfully retrieved.
  • Generation: check that the answer sticks to what was retrieved. Ragas calls this faithfulness: how factually consistent the response is with the retrieved context, on a 0 to 1 scale.

Build a small reference set first: 30 to 50 real questions, with the expected answer and the document it should come from. It becomes your regression suite every time someone changes the prompt, the model or the chunking.

Which security tests come from the OWASP LLM Top 10?

The OWASP Top 10 for LLM Applications (2025) is the most used public checklist. Several entries translate directly into test cases:

  • LLM01 Prompt Injection: user input that alters the model’s behaviour. Test it directly (“ignore your instructions and…”) and indirectly, by putting hidden instructions in a document or web page the app will read.
  • LLM02 Sensitive Information Disclosure: try to obtain another user’s data, keys or internal information.
  • LLM05 Improper Output Handling: check what happens when the model’s output is rendered or passed on, for example HTML or script in an answer displayed in the page.
  • LLM06 Excessive Agency: if the bot can call tools (send an email, cancel an order), check it cannot do more than the user is allowed to.
  • LLM07 System Prompt Leakage: ask for the instructions, in several languages and phrasings.
  • LLM09 Misinformation: questions with no answer in the documents; the right behaviour is to say so.
  • LLM10 Unbounded Consumption: very long inputs, rapid repeated requests; check limits and costs.

What does the AI Act add for a chatbot?

A test case that is easy to forget. According to the European Commission, the transparency obligations of Article 50 apply since 2 August 2026: among other things, people must be informed that they are interacting with an AI system. Check that the disclosure is there, visible, and in the user’s language. (The high-risk rules are a separate matter: since the AI Omnibus they apply from 2 December 2027.)

Which tools help?

  • promptfoo: an open-source CLI and library for evaluating and red-teaming LLM apps, usable in CI (a GitHub Action exists). Good for running your reference set and attack prompts on every change.
  • Ragas: RAG metrics such as context precision, context recall, faithfulness and response relevancy.
  • Playwright: for the chat interface itself (streaming, errors, disclosure banner, accessibility), as for any web UI.

Tools score; they do not decide. A human still reviews the failing cases and the thresholds.

A starter checklist

  1. A reference set of 30 to 50 real questions with expected facts and source documents.
  2. Each case run several times, with a target pass rate.
  3. Retrieval checked on its own (right passages returned).
  4. Answers checked against the retrieved context (no invented facts).
  5. Out-of-scope and unanswerable questions: the bot says it does not know.
  6. Direct and indirect prompt injection attempts.
  7. System prompt and other users’ data cannot be extracted.
  8. Tool calls limited to what the user is allowed to do.
  9. Model output safely rendered in the page.
  10. AI disclosure visible to the user.

Frequently asked questions

Can you write automated tests for a chatbot?

Yes, but assert properties instead of exact text: required facts present, forbidden content absent, refusal when out of scope, expected format. Run each case several times and track a pass rate.

What is the difference between context recall and faithfulness?

In Ragas, context recall measures how many of the relevant documents were retrieved; faithfulness measures how factually consistent the answer is with the retrieved context. The first tests retrieval, the second tests generation.

Must a chatbot say it is an AI under the AI Act?

According to the European Commission, the Article 50 transparency obligations apply since 2 August 2026 and include informing people that they are interacting with an AI system.

Written with AI assistance; definitions and dates were checked against the sources listed above on 5 October 2026.

Learn AI-assisted test automation

Our online programme covers Playwright, API testing, CI and AI-assisted testing, with ISTQB exam preparation. Courses are taught in French.

View our courses

Related articles