This vignette is a maintainer map for llmshieldr. It
explains where behavior lives, how data moves through the package, and
what needs to change when a contributor adds a detector, policy control,
provider adapter, or public function.
Repository Map
llmshieldr/
├── R/ package implementation and roxygen documentation
├── man/ generated Rd files; do not edit by hand
├── tests/testthat/ unit and regression tests
├── vignettes/ long-form user and developer guides
├── inst/extdata/ packaged example and security-evaluation data
├── inst/scripts/ opt-in benchmark scripts
├── DESCRIPTION dependencies, package metadata, R version
├── NAMESPACE generated exports and S3 registration
├── NEWS.md user-facing development and release notes
├── README.Rmd README source
├── README.md generated repository landing page
└── _pkgdown.yml reference and article navigation
Edit files under R/, tests, vignettes, and README
sources. Generate man/ and NAMESPACE with
roxygen2 after changing exported interfaces.
Runtime Layers
The implementation is split by responsibility:
| Layer | Main files |
|---|---|
| Core objects and findings | R/rules.R |
| Built-in policies |
R/policy.R, R/policy_controls.R
|
| Prompt and output scanning |
R/scan_prompt.R, R/scan_output.R
|
| Retrieved context |
R/scan_context.R, R/context_policy.R
|
| Conversation and documents |
R/conversation.R, R/documents.R
|
| Tools |
R/tool_guardrails.R, R/tool_policy.R
|
| Streaming | R/streaming.R |
| URL and output boundaries |
R/url_policy.R, R/output_contract.R
|
| Grounding | R/grounding.R |
| Optional scanners |
R/scanner_options.R, R/providers.R,
R/secrets.R
|
| External adapters |
R/adapters.R, R/http.R
|
| Provider orchestration | R/secure_chat.R |
| Compatibility wrappers |
R/shield_gemini.R, R/shield_ollama.R
|
| Resource and trust controls |
R/rate_guard.R, R/trust_boundary.R
|
| Audit and telemetry |
R/audit.R, R/telemetry.R,
R/stats.R
|
| Evaluation and taxonomy |
R/evaluation.R, R/owasp.R
|
| User explanations | R/explain.R |
Execution Model
The high-level guarded flow is:
prompt
│
├─ scan_prompt()
│
├─ scan_context() and context_policy()
│
├─ rate_guard() reservation
│
├─ assistant callback or ellmer provider
│
├─ scan_output()
│
├─ output_contract()
│
├─ grounding_policy()
│
├─ rate-guard update or rollback
│
└─ shieldr_result + shieldr_audit + telemetry events
secure_chat() owns this orchestration. Individual
scan_*() functions remain independently useful when an
application owns the model call or needs checks at one boundary.
Tool calls form another path:
model tool request
├─ scan_tool_call()
├─ tool allowlist, schema, authorization, spend, and loop checks
├─ dispatcher
└─ scan_tool_output()
Public Object Contracts
The package uses inspectable S3 objects.
shieldr_rule
One rule contains an ID, one regex or function, an OWASP category,
severity, action, description, and applicable stages. Exactly one of
pattern and fn must be present.
shieldr_policy
A policy contains a name, rule list, thresholds, optional rate guard,
optional trusted sources, orchestration controls, and version. Built-in
policy names are resolved by policy(); application-owned
policies are created with build_policy() or
shieldr_policy().
shieldr_report
A scanner report contains the resolved action, cleaned text, findings, risk score, policy name, check mode, timestamp, optional token count, and surface-specific metadata.
shieldr_audit and shieldr_result
An audit collects input, context, tool, and output reports plus timing, token, network, policy-version, and decision metadata. A result combines released output, the audit, an OWASP risk summary, and the final orchestration action.
Keep these shapes backward compatible. When a new field is needed, prefer an optional field with a safe default and update audit scrubbing and persistence at the same time.
Finding Schema
Detectors converge on a common finding structure:
| Field | Purpose |
|---|---|
rule_id |
Stable machine-readable identifier |
owasp |
Lowercase category such as llm02
|
severity |
low, medium, high, or critical |
action |
allow, redact, or block |
description |
Human-readable reason |
match |
Optional matched evidence |
start, end
|
Optional one-based character span |
source |
Rule, scanner, recognizer, provider, or reviewer |
confidence |
Optional calibrated detector value |
Content fields must be removed by metadata-only audit construction. New evidence-bearing fields require the same privacy review.
Adding a Deterministic Rule
- Choose a stable OWASP-prefixed ID.
- Limit the rule to relevant stages.
- Choose regex matching when exact spans are useful.
- Choose a function when structured logic is clearer.
- Add a risky positive case.
- Add a close benign case that guards against overblocking.
- Test redaction spans when the action is redact.
- Add a release note when built-in behavior changes.
new_rule <- shieldr_rule(
id = "llm02.project.identifier",
pattern = "\\bPROJECT-[0-9]{6}\\b",
owasp = "llm02",
severity = "medium",
action = "redact",
description = "Internal project identifier.",
stages = c("prompt", "context", "output")
)Rule changes require detection and overblocking evidence. A broad expression that catches the reported example but blocks ordinary domain text is incomplete.
Adding an Entity Recognizer
An entity recognizer receives one text string and returns one-based
start and end offsets. Optional output fields
include confidence, entity type, and normalized value.
recognizer <- entity_recognizer(
id = "project/member-id",
entity_type = "MEMBER_ID",
version = "1.0",
recognize = function(text) {
hit <- regexpr("\\bMEM-[0-9]{4}\\b", text, perl = TRUE)
if (hit[[1]] < 0) {
return(data.frame())
}
data.frame(
start = as.integer(hit[[1]]),
end = as.integer(hit[[1]] + attr(hit, "match.length") - 1L)
)
}
)Validate every span before returning it. Recognizer errors are skipped with a warning, so tests should cover invalid and boundary offsets.
Adding a Guardrail Provider
guardrail_provider() is the narrow extension contract
for a local or remote detector. Its function receives text,
stage, and metadata, then returns
findings.
provider <- guardrail_provider(
id = "project/detector",
version = "1.0",
stages = c("prompt", "output"),
on_error = "block",
scan = function(text, stage, metadata) {
list()
}
)Provider work should define:
- whether content leaves the process;
- supported stages;
- version evidence;
- timeout and retry behavior;
- fail-closed or explicit fail-open behavior;
- normalized findings;
- tests with mocked responses and errors.
Keep external services and executables optional. They belong in
Suggests or behind runtime checks unless the core package
cannot function without them.
Adding a Chat Provider Path
The main path is provider neutral. secure_chat()
delegates provider construction to ellmer::chat(). New
provider-specific wrappers should be avoided unless they add a distinct
safety boundary. Use:
secure_chat(
prompt = "Approved prompt.",
provider = "provider_name",
model = "model-name",
provider_args = list()
)Credentials remain the provider SDK’s responsibility. Never inspect, print, or store provider keys in reports or audits.
Tests
Tests in tests/testthat/ are organized by public surface
and regression theme: policies and rules, prompt/context/output
scanning, orchestration, reviewers, rate limits, audits, trust
boundaries, normalization, and adoption surfaces. Keep related tests
together and name new files after the behavior they verify.
Use mocks for model and HTTP behavior. The normal test suite must not require credentials, internet access, Ollama, Presidio, OPA, Gitleaks, or another external process.
Evaluation Data
inst/extdata/security_eval_cases.csv is a compact,
transparent starter corpus. Add a regression row when a behavior gap
should remain visible across policies. Keep private or large application
corpora outside the package.
The opt-in benchmark at
inst/scripts/benchmark-security-eval.R is for adoption and
release checks that do not belong in routine unit tests.
Documentation Workflow
Public functions need roxygen documentation, useful examples, and a place in the pkgdown reference index. User-facing behavior changes also need README, vignette, and NEWS updates where relevant.
Typical commands are:
devtools::document()
devtools::test()
devtools::check()
pkgdown::build_site()Do not put live API requests, credential reads, service probes,
downloads, or optional executables in evaluated examples. Keep provider
chunks eval = FALSE and document placeholder environment
variables without reading them during the build.
Compatibility and Deprecation
Use R’s standard deprecation mechanism when replacing a public path. Keep the compatibility wrapper long enough for users to migrate, name the replacement in the message, document the change, and add a regression test.
shield_gemini() and shield_ollama()
demonstrate this pattern by forwarding users to
secure_chat(provider = …).
Review Checklist
Before merging a change:
- Confirm the public contract and default behavior.
- Check failure behavior at every external boundary.
- Add positive, negative, and privacy regression tests.
- Confirm metadata-only audits remove new content fields.
- Confirm statistics do not expose prompts or credentials.
- Keep remote calls and optional binaries out of package checks.
- Review CRAN dependencies and conditional imports.
- Update generated documentation.
- Run package checks in a clean session.
- Record user-visible changes in
NEWS.md.
