Scans LLM output for sensitive data, unsafe code, agency claims, system prompt leakage, misinformation markers, and optional NLP intent signals.
Usage
scan_output(
text,
policy = "enterprise_default",
reviewer = NULL,
checks = "rules",
redaction = NULL,
scanners = scanner_options(),
contract = NULL,
show_tokens = FALSE,
stage = "output",
show_stats = FALSE
)Arguments
- text
Model output text.
- policy
A
shieldr_policyor built-in policy name such as"comprehensive".- reviewer
Optional reviewer function or object with
$chat().- checks
One of
"rules","nlp","llm", or"both".- redaction
Optional redaction strategy from
redaction_strategy().- scanners
Optional scanner configuration from
scanner_options().- contract
Optional destination contract from
output_contract().- show_tokens
Whether to attach token counts when
ellmeris available.- stage
Internal trust-boundary stage.
- show_stats
Show elapsed time, token estimate, network status, and transfer metrics when available.
Details
Output scanning is the last guardrail before model text is displayed, stored, or passed to another tool. It runs the policy rule set over the full output and adds output-specific checks for common failure modes:
fenced code blocks are scanned for unsafe code and command patterns
excessive-agency language such as "I will now" or "I have deleted"
system-prompt structural markers such as "# System" or role declarations
high-confidence medical or financial claim markers
Agency checks in checks = "rules" match known phrases and action verbs;
they cannot infer every paraphrase or distinguish every quotation from a
model's own intent. For broader intent review, supply a semantic reviewer
with checks = "both" and evaluate it on your application data.
Use checks = "nlp" when you want a lightweight local NLP-only pass over
model output. The return value is a shieldr_report() with the same scoring
and action semantics as scan_prompt().
