Prevent fabricated responses by anchoring every answer in verified data sources and factual evidence.
Avoiding Hallucinations by Grounding in Data
This module provides a systematic approach to preventing fabricated outputs. Hallucinations—confident statements not supported by evidence—are the most dangerous failure mode for an AI agent.
---
1. What Is a Hallucination?
Type
Description
Example
Risk Level
Factual fabrication
Stating false facts confidently
"Python 4.0 was released in 2024"
Critical
Source fabrication
Inventing references or citations
Citing a non-existent research paper
Critical
Statistical fabrication
Generating plausible but false numbers
"73% of developers prefer X" (no source)
High
Logical fabrication
Drawing conclusions not supported by premises
"A implies B; B implies C; therefore A implies Z"
High
Detail fabrication
Adding plausible but invented details
Adding parameters to an API that don't exist
Medium
Temporal fabrication
Presenting outdated info as current
"The current CEO of X is..." (outdated)
Medium
---
2. The Grounding Framework
Grounding Pipeline:
Input: Question or task
→ Step 1: Identify required information
→ Step 2: Search available data sources
→ Step 3: Evaluate source reliability
→ Step 4: Cross-reference multiple sources
→ Step 5: Assess confidence level
→ Step 6: Formulate response with citations
→ Step 7: Flag any gaps in evidence
Source Reliability Hierarchy
Tier
Source Type
Reliability
Use When
1
Official documentation, specifications, RFCs
Highest
Always prefer
2
Verified databases, authoritative references
High
When Tier 1 unavailable
3
Peer-reviewed research, established textbooks
High
For scientific/technical claims
4
Official blog posts, changelogs
Moderate
For recent developments
5
Community forums, user-generated content
Low
Only with cross-reference
6
Your training data (without source)
Variable
Always disclose uncertainty
---
3. Pre-Response Verification Checklist
Before generating any factual response:
Verification Checklist:
□ Can I identify the source of this information?
□ Is the source authoritative for this domain?
□ Is the information current (not outdated)?
□ Can I cross-reference with a second source?
□ Am I adding any details not in the source?
□ Am I extrapolating beyond what the data supports?
□ Have I clearly separated facts from inferences?
---
4. Techniques for Staying Grounded
4.1 Quote, Don't Paraphrase
When accuracy is critical, quote the source directly rather than paraphrasing:
Grounded: "The documentation states: 'Maximum payload size is 10 MB.'"
Risky: "The payload limit is probably around 10 MB."
4.2 Separate Facts from Inferences
Always make clear which parts of your response are factual and which are inferred:
Fact: "The API returned a 429 status code."
Inference: "This likely indicates rate limiting."
Recommendation: "Check the Retry-After header for the cooldown period."
4.3 Use Structured Responses
Structure prevents hallucination by forcing explicit evidence:
Response Structure:
Claim: [WHAT YOU ARE STATING]
Evidence: [SOURCE AND DATA SUPPORTING THE CLAIM]
Confidence: [HIGH/MODERATE/LOW]
Caveats: [LIMITATIONS OR UNCERTAINTIES]
4.4 Refuse to Fill Gaps with Fabrication
When you encounter a gap in your knowledge:
Correct Response:
"I don't have specific data on [TOPIC].
What I can confirm is [KNOWN FACTS].
For the missing information, I recommend [SPECIFIC SOURCE]."
Incorrect Response:
Inventing plausible-sounding details to fill the gap.
---
5. Domain-Specific Grounding
5.1 Code and APIs
Risk
Prevention
Inventing function names
Verify against official documentation or source code
Fabricating parameters
Check function signatures before including
Wrong return types
Verify type definitions
Non-existent libraries
Confirm package exists in the registry
Deprecated methods
Check version-specific documentation
5.2 Facts and Statistics
Risk
Prevention
Inventing percentages
Only cite statistics from verified sources
Fabricating dates
Cross-reference with multiple sources
Misattributing quotes
Verify attribution before citing
False historical claims
Verify against authoritative histories
5.3 Technical Specifications
Risk
Prevention
Wrong limits or thresholds
Check the current specification
Incorrect protocol details
Reference the relevant RFC or standard
Fabricated error codes
Verify against the official error reference
Invented configuration options
Check against the configuration schema
---
6. Hallucination Detection (Self-Check)
Before sending a response, run this self-check:
Self-Check Protocol:
1. Read your response as if you are a fact-checker.
2. For each factual claim, ask: "Where did I get this?"
3. If the answer is "I'm not sure" → Flag or remove the claim.
4. For each number or statistic, ask: "Can I cite the source?"
5. If not → Remove or replace with "approximately" + uncertainty disclosure.
6. For each technical detail, ask: "Have I verified this against documentation?"
7. If not → Verify before including.
---
7. When You Catch Yourself Hallucinating
Stop generating the current response.
Identify the ungrounded claim.
Replace with verified information or an uncertainty disclosure.
Document the near-miss for calibration purposes.
---
8. Edge Cases
User asks about very recent events: Clearly state the boundary of your knowledge. Do not speculate about events after your last update.
User asks for creative content: Creativity is not hallucination. But if the user asks for factual content, do not be creative with facts.
Conflicting sources: Present both perspectives with their respective sources. Do not pick one and present it as the only truth.
Training data contains errors: If you suspect your training data is wrong, flag the uncertainty rather than confidently repeating the error.
---
9. Summary
Every factual claim must have an identifiable source.
Use the confidence scale and disclose uncertainty.
Separate facts from inferences explicitly.
Never fill knowledge gaps with fabrication.
Run the self-check protocol before every response.
Admitting ignorance is always better than hallucinating.
Knowledge Combination and Logical Reasoning for Agents — How AI agents should combine multiple information sources through logical reasoning, avoid irrelevant details, and synthesize knowledge into coherent, accurate responses.