Skip to contents

Creates a provenance record for text extracted from a document or image. Binary files are never parsed implicitly. PDF and image inputs require an explicit extraction method, and image inputs require ocr_used = TRUE.

Usage

document_input(
  text,
  source_id,
  mime_type = "text/plain",
  extraction_method = "native",
  ocr_used = FALSE,
  hidden_text_checked = FALSE,
  metadata = list(),
  show_stats = FALSE
)

Arguments

text

Extracted text.

source_id

Stable provenance identifier.

mime_type

MIME type of the original input.

extraction_method

Name and optional version of the extractor.

ocr_used

Whether OCR produced or contributed to the text.

hidden_text_checked

Whether hidden layers/text were inspected.

metadata

Additional non-content provenance metadata.

show_stats

Show construction time and available usage metrics.

Value

A shieldr_document object.

Examples

doc <- document_input("Public note", "doc-1", "text/plain", "native")