# Chunk<no value>

{{< callout context="note" icon="outline/info-circle" >}}
One input row can leave as many chunk rows. By default the split runs locally with the gpt-4o tokenizer, and there is no request to a model. Pass another tokenizer, or replace the chunker, when you want a different split.
{{< /callout >}}

## Overview

- **Transformation Type**: Non-blocking
- **Execution Mode**: One row in, one or more pieces out
- **Package**: `ETLBox.AI`

`ChunkTransformation<TInput>` writes `Chunk<TInput>`. Each output row is one piece of the input row:

- `Source` is that input row. It is the same instance that arrived, and the transformation does not change it, so a later step can still read the id, the name, or any other field.
- `Text` is the text of this piece.
- `Context` is a heading path or other note the chunker kept with this piece. It stays empty when the chunker does not set one.
- `Index` numbers the pieces of this one row, starting at 0. The next input row starts again at 0.

A [Text Document](../text-document/) source is a common way in: one file, response, or blob becomes one row, and this transformation cuts that row. A [Markdown](../markdown/) source is the same shape with the markdown already parsed, so a header chunker can follow the headings. An [HTML](../html/) source converts HTML to that same document. A [PDF](../pdf/) source reads one PDF as an `IngestionDocument` whose sections are the pages. The usual next step is [Embedding](../embedding/), one vector per piece. The sections below are the three decisions that produce these rows: which text is split, how the split is measured, and which shape leaves the component.

## Where the text comes from

Each row has one text. You name it with `TextSelector` or with `DocumentSelector`. The two cannot be set together, because they are two different inputs: a string you build, or a document that already has sections.

### A string built from the row

`TextSelector` builds that string. Use it when the text is one field, or when several fields belong in the same piece, such as a name followed by its description.

```csharp
public class Product {
    public int Id { get; set; }
    public string Name { get; set; }
    public string Description { get; set; }
}

var source = new MemorySource<Product>();
source.DataAsList.Add(new Product {
    Id = 101,
    Name = "Trail Bottle",
    Description = "Leakproof bottle for hiking. Fits in a side pocket."
});

var chunk = new ChunkTransformation<Product> {
    TextSelector = product => $"{product.Name}. {product.Description}",
    MaxTokensPerChunk = 12,
    OverlapTokens = 2
};
var dest = new MemoryDestination<Chunk<Product>>();

source.LinkTo(chunk).LinkTo(dest);
Network.Execute(source);
```

`MaxTokensPerChunk` and `OverlapTokens` belong to the default split, described below. The small numbers in this example only force a short product text to break into more than one piece.

### A string property on the class

When one string property already is the text, mark it with `[ChunkText]` and leave both selectors empty. The transformation then reads that property. This is the same input as `TextSelector`, only the class carries the choice instead of the data flow.

```csharp
public class CatalogItem {
    public int Id { get; set; }
    public string Name { get; set; }

    [ChunkText]
    public string Description { get; set; }
}

var chunk = new ChunkTransformation<CatalogItem> {
    MaxTokensPerChunk = 8,
    OverlapTokens = 1
};
```

The attribute belongs on one readable `string`. A second `[ChunkText]`, or a marked property that cannot be read as a string, is reported when the network starts. A `TextSelector` set on the same component is the string that is split. The attribute is the fallback for when you did not set a selector.

### A dynamic row

The non-generic `ChunkTransformation` writes `Chunk<ExpandoObject>`. There is no class to mark. A property named `Document` that holds an `IngestionDocument` is the document that is split. [Markdown](../markdown/), [HTML](../html/), and [PDF](../pdf/) write that property, so the dynamic row can be linked straight through. When `Document` is absent, the transformation reads the property `Text`. `DocumentPropertyName` and `TextPropertyName` are those names. They stay `"Document"` and `"Text"` until you change them, and both are ignored once `TextSelector` or `DocumentSelector` is set.

```csharp
var source = new MemorySource();
dynamic row = new ExpandoObject();
row.Id = 301;
row.Text = "Running Shoes. Light shoes for daily runs.";
source.DataAsList.Add(row);

var chunk = new ChunkTransformation {
    MaxTokensPerChunk = 10,
    OverlapTokens = 0
};

source.LinkTo(chunk).LinkTo(new MemoryDestination());
Network.Execute(source);
```

### A document that already has structure

`DocumentSelector` is the other input. The row already holds an `IngestionDocument`, and that document is passed to the chunker unchanged, including its sections. Use this when flattening the document into one string would throw away a structure you still want the split to follow.

```csharp
public class ProductDocument {
    public int Id { get; set; }
    public IngestionDocument Document { get; set; }
}

var chunk = new ChunkTransformation<ProductDocument> {
    DocumentSelector = product => product.Document,
    MaxTokensPerChunk = 8,
    OverlapTokens = 1
};
```

`IngestionDocument` is in `Microsoft.Extensions.DataIngestion`. With a `DocumentSelector` set, `[ChunkText]` and `TextPropertyName` are not used. When the row itself is an `IngestionDocument`, that row is the document and both selectors stay empty. `[ChunkDocument]` on one readable `IngestionDocument` property is the same choice as `DocumentSelector`, carried by the class.

## How the pieces are cut

Once the text is known, a chunker cuts it. Leave `Chunker` empty and ETLBox builds a `DocumentTokenChunker`. That chunker starts a new piece when the current one reaches `MaxTokensPerChunk`, and it copies `OverlapTokens` from the end of one piece onto the start of the next. The overlap is there so a sentence that sits on the cut remains readable in both pieces. The defaults are 2000 and 500 tokens.

The length is measured by `Tokenizer`. The default is the local gpt-4o tokenizer, `TiktokenTokenizer.CreateForModel("gpt-4o")`, so the count follows that model family and nothing is sent to an API. Set `Tokenizer` when the embedding or chat model you use afterwards counts tokens differently. The pieces then follow that count. Because the unit is a token, a cut can fall inside a word.

`Tokenizer`, `MaxTokensPerChunk`, and `OverlapTokens` are the settings of this default chunker. They are read only while `Chunker` is empty.

Set `Chunker` when the text should be split by a different rule. A `SectionChunker` keeps each section of an `IngestionDocument` together. [PDF](../pdf/) makes each page a section, so that chunker keeps a page together. A `HeaderChunker` splits on the headings inside the document and stores the heading path in `Context`, for example `Guide > Install`. [Markdown](../markdown/) and [HTML](../html/) are the sources that produce that document. A chunker you assign is configured through its own options. The three properties above do not apply to it.

```csharp
var tokenizer = TiktokenTokenizer.CreateForModel("gpt-4o");

var chunk = new ChunkTransformation<ProductDocument> {
    DocumentSelector = product => product.Document,
    Chunker = new SectionChunker(new IngestionChunkerOptions(tokenizer) {
        MaxTokensPerChunk = 200,
        OverlapTokens = 0
    })
};
```

`DocumentTokenChunker` and `SectionChunker` are in `Microsoft.Extensions.DataIngestion.Chunkers`. `HeaderChunker` and `IngestionChunkerOptions` are in `Microsoft.Extensions.DataIngestion`. `TiktokenTokenizer` is in `Microsoft.ML.Tokenizers`. Any other `IngestionChunker<string>` can be assigned to `Chunker` the same way.

## The row that continues

The one-type form, `ChunkTransformation<Product>`, can emit `Chunk<Product>` directly, so `CreateOutput` stays empty. A later component often wants ordinary fields instead, for example a product id next to the piece of text. `ChunkTransformation<TInput, TOutput>` is that case. `CreateOutput` receives the `Chunk<TInput>` and returns your row. It is required for an output type other than `Chunk<TInput>`, because the component has no other way to build that type.

```csharp
public class ProductChunk {
    public int ProductId { get; set; }
    public int ChunkIndex { get; set; }
    public string ChunkText { get; set; }
}

var chunk = new ChunkTransformation<Product, ProductChunk> {
    TextSelector = product => $"{product.Name}. {product.Description}",
    CreateOutput = piece => new ProductChunk {
        ProductId = piece.Source.Id,
        ChunkIndex = piece.Index,
        ChunkText = piece.Text
    },
    MaxTokensPerChunk = 8,
    OverlapTokens = 0
};
```

`ProductChunk` is the row you pass to [Embedding](../embedding/) when each piece needs a vector.

## When a row cannot be split

A missing text, both selectors at once, or an invalid `[ChunkText]` is reported when the network starts. The flow has no single text to cut, so it does not begin. A failure on one row, such as a document that is null, follows the normal ETLBox error path. Link an error destination if that row should be kept aside and the flow should continue:

```csharp
chunk.LinkErrorTo(errorDest);
```

Further examples are in the [chunk recipes](/recipes/ai/chunk/).
