PDF Source
This article contains example code that shows the usage of the PdfSource component.
PdfSource reads one PDF resource as one IngestionDocument. A file, a web response, or a blob is one element. A folder is one element per file. Each page becomes a section, and each text block on that page becomes a paragraph. The non-generic source writes Document and StreamMetaData on an ExpandoObject.
The examples below read guide.pdf. The file has two pages.
Page 1:
Install ETLBox
Run dotnet add package ETLBox.AI.Page 2:
Other
Unrelated page body.Read one file
The whole file is one dynamic row. Document is the parsed PDF. StreamMetaData carries the path and the resource count. Sections are the pages.
string sourceFile = "guide.pdf";
var source = new PdfSource(sourceFile);
var dest = new MemoryDestination();
source.LinkTo(dest);
Network.Execute(source);
foreach (dynamic row in dest.Data) {
StreamMetaData metadata = row.StreamMetaData;
IngestionDocument parsed = row.Document;
Console.WriteLine($"RequestUri:{metadata.RequestUri}");
Console.WriteLine($"Identifier:{parsed.Identifier}");
Console.WriteLine($"Pages:{parsed.Sections.Count}");
Console.WriteLine($"Page 1:{parsed.Sections[0].Elements[0].Text}");
}
//Outputs
//RequestUri:guide.pdf
//Identifier:guide.pdf
//Pages:2
//Page 1:Install ETLBox
//Run dotnet add package ETLBox.AI.Build the row yourself
ResultSelector gets the parsed document and the StreamMetaData. Fields such as Source are set here.
public class Article
{
public string Source { get; set; }
public IngestionDocument Document { get; set; }
}
string sourceFile = "guide.pdf";
var source = new PdfSource<Article>(sourceFile) {
ResultSelector = (document, metadata) => new Article {
Source = metadata.RequestUri,
Document = document
}
};
var dest = new MemoryDestination<Article>();
source.LinkTo(dest);
Network.Execute(source);
foreach (Article row in dest.Data)
Console.WriteLine($"Source:{row.Source} Pages:{row.Document.Sections.Count}");
//Outputs
//Source:guide.pdf Pages:2Read the document itself
PdfSource<IngestionDocument> emits the parsed document. The row has no separate property for the path, because Identifier already is that URI. PageNumber on each section is the page.
string sourceFile = "guide.pdf";
var source = new PdfSource<IngestionDocument>(sourceFile);
var dest = new MemoryDestination<IngestionDocument>();
source.LinkTo(dest);
Network.Execute(source);
foreach (IngestionDocument document in dest.Data)
Console.WriteLine($"Identifier:{document.Identifier} Pages:{document.Sections.Count}");
//Outputs
//Identifier:guide.pdf Pages:2Split your own type
[ChunkDocument] names the document property. ResultSelector still builds the row, because the source does not fill a custom type on its own. With Chunker left empty, MaxTokensPerChunk and OverlapTokens cut the text.
public class Guide
{
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
string sourceFile = "guide.pdf";
var source = new PdfSource<Guide>(sourceFile) {
ResultSelector = (document, metadata) => new Guide {
Source = metadata.RequestUri,
Document = document
}
};
//No DocumentSelector needed - the attribute defines the document property
var chunk = new ChunkTransformation<Guide> {
MaxTokensPerChunk = 100,
OverlapTokens = 0
};
var dest = new MemoryDestination<Chunk<Guide>>();
source.LinkTo(chunk);
chunk.LinkTo(dest);
Network.Execute(source);
foreach (Chunk<Guide> row in dest.Data)
Console.WriteLine($"Source:{row.Source.Source} Text:{row.Text}");
//Outputs
//Source:guide.pdf Text:Install ETLBox Run dotnet add package ETLBox.AI.
//Source:guide.pdf Text:Other Unrelated page body.