HTML
HtmlSource reads one HTML resource, converts it to markdown, and parses that markdown into one IngestionDocument. Headings stay in the document, so a later chunk step can split by header and still know which section a piece came from.
One resource leaves as one document, even when the HTML is empty, and a folder is one document per file in that folder. Markdown that is already markdown remains Markdown. Plain text, one row per file, remains Text Document.
One file becomes one document
- Component: Source
- Execution: One resource in, one document out
- Package:
ETLBox.AI.Html
An HTML page ends up as the same kind of document
Markdown produces from a .md file. The resource is read to the end, converted from HTML to markdown, and then parsed with MarkdownReader. Headings in the page stay headings in the document, which is what a later header chunker follows. The document Identifier is the resource URI. An empty resource is still one row.
Uri, Folder, ResourceType, encoding, and RowModificationAction work as they do on the other streaming connectors. Files, HTTP, and Azure Blob Storage are described in
Shared Features. HtmlSource<TOutput> can write your type, the IngestionDocument itself, or an ExpandoObject.
Install ETLBox.AI.Html. That package reads the HTML and brings in ETLBox.AI, which holds
chunking, chat, and embeddings. The conversion itself is the next section. After that, the document leaves as the document, on a class, or on a dynamic row, and the last step splits it on the headings.
How the HTML is converted
The conversion uses
ReverseMarkdown, and the markdown that comes out is what MarkdownReader parses. ConverterConfig is the ReverseMarkdown.Config for that first step. A new source removes script and style elements before the conversion, so a script stays out of the text and the headings are what remain.
Edit that object when you want further ReverseMarkdown options and still want those elements removed. A config you assign yourself starts from ReverseMarkdown’s defaults. Put the script and style removal on that object again:
var config = new ReverseMarkdown.Config();
config.Preprocess.RemoveScripts().RemoveStyles();
var source = new HtmlSource("guide.html") {
ConverterConfig = config
};Use the parsed document as the row
From here the row looks like the one from
Markdown. The simplest form is the document itself. HtmlSource<IngestionDocument> emits the parse. Identifier is the path, the URL, or the blob name. Sections hold the elements the reader produced, including the headings.
var source = new HtmlSource<IngestionDocument>("guide.html");
var dest = new MemoryDestination<IngestionDocument>();
source.LinkTo(dest);
Network.Execute(source);
foreach (IngestionDocument document in dest.Data) {
Console.WriteLine(document.Identifier);
Console.WriteLine(document.Sections[0].Elements[0].Text);
}Chunk treats an IngestionDocument row as the document, so DocumentSelector stays empty. The heading split further down uses this form.
Map the document onto your type
When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. ResultSelector receives the parsed document and the StreamMetaData and returns the row. Properties you leave unset stay at their default.
Mark the document property with [ChunkDocument] when the chunk step below should find it. The attribute belongs on one readable IngestionDocument. The source does not read the attribute and does not fill the property. ResultSelector still builds the row, and
Chunk reads the marked property when the network starts. DocumentSelector then stays empty. A DocumentSelector set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.
public class Article {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new HtmlSource<Article>("guide.html") {
ResultSelector = (document, metadata) => new Article {
Source = metadata.RequestUri,
Document = document
}
};A second [ChunkDocument], or a marked property that is not a readable IngestionDocument, is reported when the network starts. [ChunkDocument] and [ChunkText] belong on different types, because one names a document and the other names a string.
Read the file without a class
With no class to mark, the non-generic HtmlSource writes an ExpandoObject. The row has Document for the parsed HTML and StreamMetaData for the path, the resource count, and the rest of the streaming metadata. Document is already the property
Chunk reads when it holds an IngestionDocument, so this row links straight through to the split below.
var source = new HtmlSource("guide.html");
var dest = new MemoryDestination();
source.LinkTo(dest);
Network.Execute(source);
foreach (dynamic row in dest.Data) {
StreamMetaData metadata = row.StreamMetaData;
IngestionDocument parsed = row.Document;
Console.WriteLine(metadata.RequestUri);
Console.WriteLine(parsed.Identifier);
Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}ResultSelector replaces that pair when the dynamic row should look different, in the same way it builds a class above.
Read a folder, a URL, or a blob
Each of those row shapes is one element per resource. Uri names one resource. Folder names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. RequestCount follows the sequence.
var source = new HtmlSource {
Folder = "pages"
};ResourceType selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. GetNextUri and HasNextUri walk a list of addresses when the resources are neither a single path nor a folder. SkipRows drops leading lines of the HTML, and the remainder of that resource is converted and parsed as the one document.
Split the HTML by heading
This is the step the headings were kept for. After the conversion they are markdown headings, so the split is the same one
Markdown uses. Link the dynamic row, or the IngestionDocument, straight to
Chunk. A typed row points that transformation at the document with [ChunkDocument] or DocumentSelector, as set up above.
HeaderChunker splits on those headings and keeps the heading path. Context on each chunk is that path, such as Guide > Install. Text is the body under those headings. The limits belong on IngestionChunkerOptions, because a chunker you assign is configured through its own options. MaxTokensPerChunk and OverlapTokens on the transformation apply while Chunker is empty.
var source = new HtmlSource<IngestionDocument>("guide.html");
var chunk = new ChunkTransformation<IngestionDocument> {
Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
MaxTokensPerChunk = 512,
OverlapTokens = 50
})
};
source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);HeaderChunker, IngestionChunkerOptions, and IngestionDocument are in Microsoft.Extensions.DataIngestion. TiktokenTokenizer is in Microsoft.ML.Tokenizers.
On the class from above, [ChunkDocument] already names the property, so DocumentSelector stays empty:
public class Guide {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new HtmlSource<Guide>("guide.html") {
ResultSelector = (document, metadata) => new Guide {
Source = metadata.RequestUri,
Document = document
}
};
var chunk = new ChunkTransformation<Guide> {
Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
MaxTokensPerChunk = 512,
OverlapTokens = 50
})
};When reading a resource fails
A custom type without ResultSelector fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or converting one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:
source.LinkErrorTo(errorDest);Further examples are in the HTML recipes.