Creates a provenance record for text extracted from a document or image.
Binary files are never parsed implicitly. PDF and image inputs require an
explicit extraction method, and image inputs require ocr_used = TRUE.
Usage
document_input(
text,
source_id,
mime_type = "text/plain",
extraction_method = "native",
ocr_used = FALSE,
hidden_text_checked = FALSE,
metadata = list(),
show_stats = FALSE
)Arguments
- text
Extracted text.
- source_id
Stable provenance identifier.
- mime_type
MIME type of the original input.
- extraction_method
Name and optional version of the extractor.
- ocr_used
Whether OCR produced or contributed to the text.
Whether hidden layers/text were inspected.
- metadata
Additional non-content provenance metadata.
- show_stats
Show construction time and available usage metrics.
