Markdown
TextDocumentSource reads a file as one string. MarkdownSource reads the same kind of resource and parses the markdown into one IngestionDocument. Headings stay in that document, so a later chunk step can split by header and still know which section a piece came from.
One resource leaves as one document, even when the markdown is empty, and a folder is one document per file in that folder. Plain text, one row per file, remains Text Document. HTML that should become this same document is HTML. A PDF, one document per file with a section per page, is PDF.
One file becomes one document
- Component: Source
- Execution: One resource in, one document out
- Package:
ETLBox.AI
Text Document would hand you this file as one string, headings included as characters. Here the same resource is parsed, so the headings survive as structure. Uri, Folder, ResourceType, encoding, and RowModificationAction work as they do on the other streaming connectors, and files, HTTP, and Azure Blob Storage are described in
Shared Features. The stream is read to the end and parsed with MarkdownReader from Microsoft.Extensions.DataIngestion. That parse is one IngestionDocument. The document Identifier is the resource URI. An empty file is still one row.
MarkdownSource<TOutput> can write your type, the IngestionDocument itself, or an ExpandoObject. The next sections are those three shapes. After that, the same reading covers a folder, a URL, or a blob, and the last step splits the document on the headings that were kept.
Use the parsed document as the row
The simplest row is the document. MarkdownSource<IngestionDocument> emits the parse directly. Identifier is the path, the URL, or the blob name. Sections hold the elements the reader produced, including the headings.
var source = new MarkdownSource<IngestionDocument>("guide.md");
var dest = new MemoryDestination<IngestionDocument>();
source.LinkTo(dest);
Network.Execute(source);
foreach (IngestionDocument document in dest.Data) {
Console.WriteLine(document.Identifier);
Console.WriteLine(document.Sections[0].Elements.Count());
}Chunk treats an IngestionDocument row as the document, so DocumentSelector stays empty. The heading split further down uses this form.
Map the document onto your type
When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. ResultSelector receives the parsed document and the StreamMetaData and returns the row. Properties you leave unset stay at their default.
Mark the document property with [ChunkDocument] when the chunk step below should find it. The attribute belongs on one readable IngestionDocument. The source does not read the attribute and does not fill the property. ResultSelector still builds the row, and
Chunk reads the marked property when the network starts. DocumentSelector then stays empty. A DocumentSelector set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.
public class Article {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new MarkdownSource<Article>("guide.md") {
ResultSelector = (document, metadata) => new Article {
Source = metadata.RequestUri,
Document = document
}
};A second [ChunkDocument], or a marked property that is not a readable IngestionDocument, is reported when the network starts. [ChunkDocument] and [ChunkText] belong on different types, because one names a document and the other names a string.
Read the file without a class
With no class to mark, the non-generic MarkdownSource writes an ExpandoObject. The row has Document for the parsed markdown and StreamMetaData for the path, the resource count, and the rest of the streaming metadata. Document is already the property
Chunk reads when it holds an IngestionDocument, so this row links straight through to the split below.
var source = new MarkdownSource("guide.md");
var dest = new MemoryDestination();
source.LinkTo(dest);
Network.Execute(source);
foreach (dynamic row in dest.Data) {
StreamMetaData metadata = row.StreamMetaData;
IngestionDocument parsed = row.Document;
Console.WriteLine(metadata.RequestUri);
Console.WriteLine(parsed.Identifier);
Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}ResultSelector replaces that pair when the dynamic row should look different, in the same way it builds a class above.
Read a folder, a URL, or a blob
Each of those row shapes is one element per resource. Uri names one resource. Folder names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. RequestCount follows the sequence.
var source = new MarkdownSource {
Folder = "guides"
};ResourceType selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. GetNextUri and HasNextUri walk a list of addresses when the resources are neither a single path nor a folder. SkipRows drops leading lines, and the remainder of that resource is parsed as the one document.
Split the markdown by heading
This is the step the headings were kept for. Link the dynamic row, or the IngestionDocument, straight to
Chunk. A typed row points that transformation at the document with [ChunkDocument] or DocumentSelector, as set up above.
HeaderChunker splits on the headings and keeps the heading path. Context on each chunk is that path, such as Guide > Install. Text is the body under those headings. The limits belong on IngestionChunkerOptions, because a chunker you assign is configured through its own options. MaxTokensPerChunk and OverlapTokens on the transformation apply while Chunker is empty.
var source = new MarkdownSource<IngestionDocument>("guide.md");
var chunk = new ChunkTransformation<IngestionDocument> {
Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
MaxTokensPerChunk = 512,
OverlapTokens = 50
})
};
source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);HeaderChunker, IngestionChunkerOptions, and IngestionDocument are in Microsoft.Extensions.DataIngestion. TiktokenTokenizer is in Microsoft.ML.Tokenizers.
On the class from above, [ChunkDocument] already names the property, so DocumentSelector stays empty:
public class Guide {
public string Source { get; set; }
[ChunkDocument]
public IngestionDocument Document { get; set; }
}
var source = new MarkdownSource<Guide>("guide.md") {
ResultSelector = (document, metadata) => new Guide {
Source = metadata.RequestUri,
Document = document
}
};
var chunk = new ChunkTransformation<Guide> {
Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
MaxTokensPerChunk = 512,
OverlapTokens = 50
})
};When reading a resource fails
A custom type without ResultSelector fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or parsing one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:
source.LinkErrorTo(errorDest);Further examples are in the markdown recipes.