ctx.pdf reads the text layer of a PDF, page by page, with provenance. This is the first PDF slice and it is deliberately narrow: it reads text and does nothing else (framework/pdf-types.ts).
#What this capability is
ctx.pdf.read(name, options) returns the document's text, page by page:
pagesis the requested pages, in ascending document order.meta.pageCountis how many pages the document actually has.sourceHashis a SHA-256 over the bytes that were read.enginenames the engine and the version that was loaded for this read.
The engine is pdfjs-dist, loaded lazily on the first read. If it cannot be loaded, the read refuses with PDF_ENGINE_UNAVAILABLE naming the module, rather than crashing at import time.
The document is named through the ctx.files jail. ctx.pdf holds no jail of its own and never touches node:fs: it reads the bytes from the files facade. An absolute host path, or a .. traversal, is refused by the jail and surfaces here as PDF_FILES_REQUIRED. A PDF outside the Robot's data directory cannot be reached at all, by construction.
#Pages and provenance
options.pages takes 'all' (the default), a range string such as '1-20' that is inclusive and 1-based, or an array of 1-based page numbers. Whatever the request order, the answer is always in ascending document order.
One read may return at most 200 pages. A selection over that cap, or one that falls outside the document's page count, is refused with PDF_PAGE_LIMIT, never clamped and never truncated. A caller that asked for page 9 of a 3-page document is holding the wrong document and the message says so.
A page's text is the engine's own text items in the engine's own order, each string appended as it stands, with a single newline where the engine reports an end of line. Nothing is re-wrapped or re-spaced, because a caller extracting an amount from a line must see the line the document carries.
const doc = await ctx.pdf.read('invoices/alpha.pdf', { pages: '1-2' });
for (const page of doc.pages) {
ctx.log.info(`page ${page.page}: ${page.text}`);
}
ctx.log.info(`document has ${doc.meta.pageCount} pages, engine ${doc.engine.name} ${doc.engine.version}`);
#Refusals
Every failure is a named SystemException:
PDF_FILES_REQUIRED, the name is empty or not readable through the files jail.PDF_ENGINE_UNAVAILABLE, the engine cannot be loaded.PDF_MALFORMED, the bytes are not a readable PDF.PDF_ENCRYPTED, the document is encrypted and no password is supplied.PDF_NO_TEXT_LAYER, a requested page carries no text at all.PDF_PAGE_LIMIT, the selection cannot be served.
The runner adds PDF_NOT_DECLARED and PDF_UNAVAILABLE when no facade was supplied at all. A page with no text layer, which is what a scan or an image-only export has, is PDF_NO_TEXT_LAYER naming the pages, never an empty string. If any requested page is empty, the whole read refuses, so a mixed document whose page 2 is a scan does not return one empty page inside an otherwise successful result.
#What is not built
There is no PDF writing of any kind: no split, no join, no export. There is no OCR, no AcroForm field extraction, no region or template extraction and no document understanding. These are absent from the type, not merely unimplemented, so an author cannot call a promise this build does not keep. OCR is what a scanned document would need and the refusal says as much.
The UiPath converter still refuses ReadPDF, ReadPDFText and ExtractPDF by name, so no converted workflow emits ctx.pdf yet. Whether ReadPDFText should now map onto ctx.pdf.read is an open operator question, recorded and not decided.
#Things that catch people out
Read a long document in ranges. 'all' is bounded by the same 200-page cap as an explicit range and a larger document refuses until you pass a range.
sourceHash covers the bytes that were read, whole, not the selected pages, so it identifies the document rather than the selection.
An empty user password on an encrypted document opens it. That is measured behaviour. A password that is genuinely needed is the PDF_ENCRYPTED refusal above.
#What state it is in
PDF slice 1 is finished, verified and in the clean-required gate. The writer, OCR, AcroForm and document-understanding families are not built.