PdfSource reads one PDF and parses it into one IngestionDocument. Each page becomes a section, and the text blocks on that page become paragraphs. A later chunk step can split those pages and still know which page a piece came from.
One resource leaves as one document, even when a page has no extractable text, and a folder is one document per file in that folder. Plain text, one row per file, remains Text Document. Markdown remains Markdown. HTML that is converted to markdown first is HTML.
One file becomes one document
- Component: Source
- Execution: One resource in, one document out
- Package:
ETLBox.AI.Pdf
Markdown and
HTML keep headings. A PDF keeps pages. The file is opened like the other streaming connectors: Uri, Folder, ResourceType, and RowModificationAction work the same way, and files, HTTP, and Azure Blob Storage are described in
Shared Features. The stream is then read to the end and parsed with
PdfPig. Each page becomes a section. Each text block on that page becomes a paragraph, in reading order. Blank blocks are left out. The document Identifier is the resource URI, and PageNumber is set on the section and on each paragraph. That page structure is what a later section chunker keeps together.
PdfSource<TOutput> can write your type, the IngestionDocument itself, or an ExpandoObject. Install ETLBox.AI.Pdf. That package reads the PDF and brings in ETLBox.AI, which holds
chunking, chat, and embeddings. The next sections are those three shapes. After that, the same reading covers a folder, a URL, or a blob, and the last step splits the document along the pages.
Use the parsed document as the row
The simplest row is the document. PdfSource<IngestionDocument> emits the parse directly. Identifier is the path, the URL, or the blob name. Sections are the pages. Elements on a section are the paragraphs of that page.
var source = new PdfSource<IngestionDocument>("guide.pdf");
var dest = new MemoryDestination<IngestionDocument>();
source.LinkTo(dest);
Network.Execute(source);
foreach (IngestionDocument document in dest.Data) {
Console.WriteLine(document.Identifier);
Console.WriteLine(document.Sections.Count);
Console.WriteLine(document.Sections[0].PageNumber);
Console.WriteLine(document.Sections[0].Elements[0].Text);
}Chunk treats an IngestionDocument row as the document, so DocumentSelector stays empty. The page split further down uses this form.
Map the document onto your type
When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. ResultSelector receives the parsed document and the StreamMetaData and returns the row. Properties you leave unset stay at their default.
Mark the document property with [ChunkDocument] when the chunk step below should find it. The attribute belongs on one readable IngestionDocument. The source does not read the attribute and does not fill the property. ResultSelector still builds the row, and
Chunk reads the marked property when the network starts. DocumentSelector then stays empty. A DocumentSelector set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.
public class Article {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new PdfSource<Article>("guide.pdf") {
ResultSelector = (document, metadata) => new Article {
Source = metadata.RequestUri,
Document = document
}
};A second [ChunkDocument], or a marked property that is not a readable IngestionDocument, is reported when the network starts. [ChunkDocument] and [ChunkText] belong on different types, because one names a document and the other names a string.
Read the file without a class
With no class to mark, the non-generic PdfSource writes an ExpandoObject. The row has Document for the parsed PDF and StreamMetaData for the path, the resource count, and the rest of the streaming metadata. Document is already the property
Chunk reads when it holds an IngestionDocument, so this row links straight through to the split below.
var source = new PdfSource("guide.pdf");
var dest = new MemoryDestination();
source.LinkTo(dest);
Network.Execute(source);
foreach (dynamic row in dest.Data) {
StreamMetaData metadata = row.StreamMetaData;
IngestionDocument parsed = row.Document;
Console.WriteLine(metadata.RequestUri);
Console.WriteLine(parsed.Identifier);
Console.WriteLine(parsed.Sections.Count);
Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}ResultSelector replaces that pair when the dynamic row should look different, in the same way it builds a class above.
Read a folder, a URL, or a blob
Each of those row shapes is one element per resource. Uri names one resource. Folder names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. RequestCount follows the sequence.
var source = new PdfSource {
Folder = "manuals"
};ResourceType selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. GetNextUri and HasNextUri walk a list of addresses when the resources are neither a single path nor a folder.
SkipRows has no lines to drop in a PDF. An assigned value is discarded, the property stays at 0, and the whole file remains the one document.
Split the PDF into chunks
The pages were kept so this step can follow them. Link the dynamic row, or the IngestionDocument, straight to
Chunk. A typed row points that transformation at the document with [ChunkDocument] or DocumentSelector, as set up above.
Leave Chunker empty and the default token chunker cuts the document when a piece reaches MaxTokensPerChunk. When a piece should stay inside one page, use SectionChunker. Each page is already a section, so that chunker keeps a page together. The limits belong on IngestionChunkerOptions, because a chunker you assign is configured through its own options. MaxTokensPerChunk and OverlapTokens on the transformation apply while Chunker is empty.
var source = new PdfSource<IngestionDocument>("guide.pdf");
var chunk = new ChunkTransformation<IngestionDocument> {
Chunker = new SectionChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
MaxTokensPerChunk = 512,
OverlapTokens = 0
})
};
source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);SectionChunker, IngestionChunkerOptions, and IngestionDocument are in Microsoft.Extensions.DataIngestion. TiktokenTokenizer is in Microsoft.ML.Tokenizers.
On the class from above, [ChunkDocument] already names the property, so DocumentSelector stays empty. With Chunker left empty, the token limits on the transformation apply:
public class Guide {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new PdfSource<Guide>("guide.pdf") {
ResultSelector = (document, metadata) => new Guide {
Source = metadata.RequestUri,
Document = document
}
};
var chunk = new ChunkTransformation<Guide> {
MaxTokensPerChunk = 100,
OverlapTokens = 0
};When reading a resource fails
A custom type without ResultSelector fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or parsing one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:
source.LinkErrorTo(errorDest);Further examples are in the PDF recipes.