One input row can leave as many chunk rows. By default the split runs locally with the gpt-4o tokenizer, and there is no request to a model. Pass another tokenizer, or replace the chunker, when you want a different split.

Overview

  • Transformation Type: Non-blocking
  • Execution Mode: One row in, one or more pieces out
  • Package: ETLBox.AI

ChunkTransformation<TInput> writes Chunk<TInput>. Each output row is one piece of the input row:

  • Source is that input row. It is the same instance that arrived, and the transformation does not change it, so a later step can still read the id, the name, or any other field.
  • Text is the text of this piece.
  • Context is a heading path or other note the chunker kept with this piece. It stays empty when the chunker does not set one.
  • Index numbers the pieces of this one row, starting at 0. The next input row starts again at 0.

A Text Document source is a common way in: one file, response, or blob becomes one row, and this transformation cuts that row. A Markdown source is the same shape with the markdown already parsed, so a header chunker can follow the headings. An HTML source converts HTML to that same document. A PDF source reads one PDF as an IngestionDocument whose sections are the pages. The usual next step is Embedding, one vector per piece. The sections below are the three decisions that produce these rows: which text is split, how the split is measured, and which shape leaves the component.

Where the text comes from

Each row has one text. You name it with TextSelector or with DocumentSelector. The two cannot be set together, because they are two different inputs: a string you build, or a document that already has sections.

A string built from the row

TextSelector builds that string. Use it when the text is one field, or when several fields belong in the same piece, such as a name followed by its description.

public class Product {
    public int Id { get; set; }
    public string Name { get; set; }
    public string Description { get; set; }
}

var source = new MemorySource<Product>();
source.DataAsList.Add(new Product {
    Id = 101,
    Name = "Trail Bottle",
    Description = "Leakproof bottle for hiking. Fits in a side pocket."
});

var chunk = new ChunkTransformation<Product> {
    TextSelector = product => $"{product.Name}. {product.Description}",
    MaxTokensPerChunk = 12,
    OverlapTokens = 2
};
var dest = new MemoryDestination<Chunk<Product>>();

source.LinkTo(chunk).LinkTo(dest);
Network.Execute(source);

MaxTokensPerChunk and OverlapTokens belong to the default split, described below. The small numbers in this example only force a short product text to break into more than one piece.

A string property on the class

When one string property already is the text, mark it with [ChunkText] and leave both selectors empty. The transformation then reads that property. This is the same input as TextSelector, only the class carries the choice instead of the data flow.

public class CatalogItem {
    public int Id { get; set; }
    public string Name { get; set; }

    [ChunkText]
    public string Description { get; set; }
}

var chunk = new ChunkTransformation<CatalogItem> {
    MaxTokensPerChunk = 8,
    OverlapTokens = 1
};

The attribute belongs on one readable string. A second [ChunkText], or a marked property that cannot be read as a string, is reported when the network starts. A TextSelector set on the same component is the string that is split. The attribute is the fallback for when you did not set a selector.

A dynamic row

The non-generic ChunkTransformation writes Chunk<ExpandoObject>. There is no class to mark. A property named Document that holds an IngestionDocument is the document that is split. Markdown, HTML, and PDF write that property, so the dynamic row can be linked straight through. When Document is absent, the transformation reads the property Text. DocumentPropertyName and TextPropertyName are those names. They stay "Document" and "Text" until you change them, and both are ignored once TextSelector or DocumentSelector is set.

var source = new MemorySource();
dynamic row = new ExpandoObject();
row.Id = 301;
row.Text = "Running Shoes. Light shoes for daily runs.";
source.DataAsList.Add(row);

var chunk = new ChunkTransformation {
    MaxTokensPerChunk = 10,
    OverlapTokens = 0
};

source.LinkTo(chunk).LinkTo(new MemoryDestination());
Network.Execute(source);

A document that already has structure

DocumentSelector is the other input. The row already holds an IngestionDocument, and that document is passed to the chunker unchanged, including its sections. Use this when flattening the document into one string would throw away a structure you still want the split to follow.

public class ProductDocument {
    public int Id { get; set; }
    public IngestionDocument Document { get; set; }
}

var chunk = new ChunkTransformation<ProductDocument> {
    DocumentSelector = product => product.Document,
    MaxTokensPerChunk = 8,
    OverlapTokens = 1
};

IngestionDocument is in Microsoft.Extensions.DataIngestion. With a DocumentSelector set, [ChunkText] and TextPropertyName are not used. When the row itself is an IngestionDocument, that row is the document and both selectors stay empty. [ChunkDocument] on one readable IngestionDocument property is the same choice as DocumentSelector, carried by the class.

How the pieces are cut

Once the text is known, a chunker cuts it. Leave Chunker empty and ETLBox builds a DocumentTokenChunker. That chunker starts a new piece when the current one reaches MaxTokensPerChunk, and it copies OverlapTokens from the end of one piece onto the start of the next. The overlap is there so a sentence that sits on the cut remains readable in both pieces. The defaults are 2000 and 500 tokens.

The length is measured by Tokenizer. The default is the local gpt-4o tokenizer, TiktokenTokenizer.CreateForModel("gpt-4o"), so the count follows that model family and nothing is sent to an API. Set Tokenizer when the embedding or chat model you use afterwards counts tokens differently. The pieces then follow that count. Because the unit is a token, a cut can fall inside a word.

Tokenizer, MaxTokensPerChunk, and OverlapTokens are the settings of this default chunker. They are read only while Chunker is empty.

Set Chunker when the text should be split by a different rule. A SectionChunker keeps each section of an IngestionDocument together. PDF makes each page a section, so that chunker keeps a page together. A HeaderChunker splits on the headings inside the document and stores the heading path in Context, for example Guide > Install. Markdown and HTML are the sources that produce that document. A chunker you assign is configured through its own options. The three properties above do not apply to it.

var tokenizer = TiktokenTokenizer.CreateForModel("gpt-4o");

var chunk = new ChunkTransformation<ProductDocument> {
    DocumentSelector = product => product.Document,
    Chunker = new SectionChunker(new IngestionChunkerOptions(tokenizer) {
        MaxTokensPerChunk = 200,
        OverlapTokens = 0
    })
};

DocumentTokenChunker and SectionChunker are in Microsoft.Extensions.DataIngestion.Chunkers. HeaderChunker and IngestionChunkerOptions are in Microsoft.Extensions.DataIngestion. TiktokenTokenizer is in Microsoft.ML.Tokenizers. Any other IngestionChunker<string> can be assigned to Chunker the same way.

The row that continues

The one-type form, ChunkTransformation<Product>, can emit Chunk<Product> directly, so CreateOutput stays empty. A later component often wants ordinary fields instead, for example a product id next to the piece of text. ChunkTransformation<TInput, TOutput> is that case. CreateOutput receives the Chunk<TInput> and returns your row. It is required for an output type other than Chunk<TInput>, because the component has no other way to build that type.

public class ProductChunk {
    public int ProductId { get; set; }
    public int ChunkIndex { get; set; }
    public string ChunkText { get; set; }
}

var chunk = new ChunkTransformation<Product, ProductChunk> {
    TextSelector = product => $"{product.Name}. {product.Description}",
    CreateOutput = piece => new ProductChunk {
        ProductId = piece.Source.Id,
        ChunkIndex = piece.Index,
        ChunkText = piece.Text
    },
    MaxTokensPerChunk = 8,
    OverlapTokens = 0
};

ProductChunk is the row you pass to Embedding when each piece needs a vector.

When a row cannot be split

A missing text, both selectors at once, or an invalid [ChunkText] is reported when the network starts. The flow has no single text to cut, so it does not begin. A failure on one row, such as a document that is null, follows the normal ETLBox error path. Link an error destination if that row should be kept aside and the flow should continue:

chunk.LinkErrorTo(errorDest);

Further examples are in the chunk recipes.