One resource leaves as one row, even when the text is empty, and a folder is one row per file in that folder. The line-by-line reader remains Text.

One resource becomes one row

  • Component: Source
  • Execution: One resource in, one row out
  • Package: ETLBox.AI

A file, a web response, or a blob is opened the same way as on the other streaming connectors. Uri, Folder, ResourceType, encoding, and RowModificationAction are described in Shared Features. What changes is the row. The stream is read through to the end, and that complete text is the one element that leaves. An empty file is still that one row.

TextDocumentSource<TOutput> can write your own type, a plain string, or an ExpandoObject. The rest of this page follows that row. First the text lands on a class, or on a dynamic row when there is no class. Then the same reading covers a folder, a URL, or a blob. At the end the text is split into chunks, and every piece still knows which resource it came from.

Map the text onto your type

Once the file is one string, that string has to land somewhere. When the class already has a place for it, two attributes tell the source where to write. The source creates the object and then sets the marked properties, so the type needs a public parameterless constructor.

  • [DocumentContent] receives the full text. One writable string is required.
  • [DocumentRequestUri] is optional. It receives RequestUri only, as a string. The rest of StreamMetaData stays on the metadata and is not copied onto the object.

Any property without a mark stays at its default. A title stays empty until you set it.

public class Article {
    public string Title { get; set; }

    [DocumentContent]
    public string Content { get; set; }

    [DocumentRequestUri]
    public string Source { get; set; }
}

var source = new TextDocumentSource<Article>("catalog.md");
var dest = new MemoryDestination<Article>();

source.LinkTo(dest);
Network.Execute(source);

The attributes cover the text and the path. When the row should also carry a title, or any other field you take from the metadata, build it yourself. ResultSelector receives the full text and the StreamMetaData and returns the row. With a selector set, the attributes stay unused, because the function already decided what the row contains.

var source = new TextDocumentSource<Article>("catalog.md") {
    ResultSelector = (text, metadata) => new Article {
        Title = "Catalog",
        Source = metadata.RequestUri,
        Content = text
    }
};

The narrowest row is the text alone. TextDocumentSource<string> emits that string, which is enough when the path is already known and only the text should enter the flow. A ResultSelector still applies if you set one, including for string.

Read the file without a class

With no class to mark, the non-generic TextDocumentSource writes an ExpandoObject. The row has Text for the full text and StreamMetaData for the path, the resource count, and the rest of the streaming metadata. Text is already the property Chunk reads, so this row can go straight into the split below.

var source = new TextDocumentSource<ExpandoObject>("readme.md");
var dest = new MemoryDestination<ExpandoObject>();

source.LinkTo(dest);
Network.Execute(source);

foreach (dynamic row in dest.Data) {
    StreamMetaData metadata = row.StreamMetaData;
    Console.WriteLine(metadata.RequestUri);
    Console.WriteLine(row.Text);
}

ResultSelector replaces that pair when the dynamic row should look different, in the same way it replaces the attributes on a class.

Read a folder, a URL, or a blob

Each of those row shapes is one element per resource. Uri names one resource. Folder names every file in that folder, and each file becomes its own row. Subfolders stay out of the read. RequestCount follows the sequence, so a later step can tell the files apart.

var source = new TextDocumentSource {
    Folder = "notes"
};

ResourceType selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. GetNextUri and HasNextUri walk a list of addresses when the resources are neither a single path nor a folder. SkipRows drops leading lines, and the remainder of that resource is the one document.

Split the text into chunks

The row is now one text with its source attached, which is what a chunk step wants. Link the dynamic row straight to Chunk. A typed row points that transformation at the property [DocumentContent] fills, with TextSelector or [ChunkText].

var source = new TextDocumentSource("readme.md");
var chunk = new ChunkTransformation {
    MaxTokensPerChunk = 8,
    OverlapTokens = 0
};

source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<ExpandoObject>>());
Network.Execute(source);

This split measures tokens in a single string. A markdown file read here keeps its headings as characters inside that string. When those headings should decide the cut, read the file with Markdown, which parses them into a document a header chunker can follow. HTML converts an HTML file into that same document. A PDF is read with PDF, one document per file, with each page as its own section.

When reading a resource fails

The network reports a class it cannot fill when it starts: no [DocumentContent], two content or URI marks, or a mark that is not a writable string. The source would otherwise have no single property for the text. A failure while reading one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:

source.LinkErrorTo(errorDest);

Further examples are in the text document recipes.