# Markdown<no value>

{{< callout context="note" icon="outline/info-circle" >}}
One resource leaves as one document, even when the markdown is empty, and a folder is one document per file in that folder. Plain text, one row per file, remains [Text Document](../text-document/). HTML that should become this same document is [HTML](../html/). A PDF, one document per file with a section per page, is [PDF](../pdf/).
{{< /callout >}}

## One file becomes one document

- **Component**: Source
- **Execution**: One resource in, one document out
- **Package**: `ETLBox.AI`

[Text Document](../text-document/) would hand you this file as one string, headings included as characters. Here the same resource is parsed, so the headings survive as structure. `Uri`, `Folder`, `ResourceType`, encoding, and `RowModificationAction` work as they do on the other streaming connectors, and files, HTTP, and Azure Blob Storage are described in [Shared Features](../streaming-connectors/shared/). The stream is read to the end and parsed with `MarkdownReader` from `Microsoft.Extensions.DataIngestion`. That parse is one `IngestionDocument`. The document `Identifier` is the resource URI. An empty file is still one row.

`MarkdownSource<TOutput>` can write your type, the `IngestionDocument` itself, or an `ExpandoObject`. The next sections are those three shapes. After that, the same reading covers a folder, a URL, or a blob, and the last step splits the document on the headings that were kept.

## Use the parsed document as the row

The simplest row is the document. `MarkdownSource<IngestionDocument>` emits the parse directly. `Identifier` is the path, the URL, or the blob name. `Sections` hold the elements the reader produced, including the headings.

```csharp
var source = new MarkdownSource<IngestionDocument>("guide.md");
var dest = new MemoryDestination<IngestionDocument>();

source.LinkTo(dest);
Network.Execute(source);

foreach (IngestionDocument document in dest.Data) {
    Console.WriteLine(document.Identifier);
    Console.WriteLine(document.Sections[0].Elements.Count());
}
```

[Chunk](../chunk/) treats an `IngestionDocument` row as the document, so `DocumentSelector` stays empty. The heading split further down uses this form.

## Map the document onto your type

When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. `ResultSelector` receives the parsed document and the `StreamMetaData` and returns the row. Properties you leave unset stay at their default.

Mark the document property with `[ChunkDocument]` when the chunk step below should find it. The attribute belongs on one readable `IngestionDocument`. The source does not read the attribute and does not fill the property. `ResultSelector` still builds the row, and [Chunk](../chunk/) reads the marked property when the network starts. `DocumentSelector` then stays empty. A `DocumentSelector` set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.

```csharp
public class Article {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new MarkdownSource<Article>("guide.md") {
    ResultSelector = (document, metadata) => new Article {
        Source = metadata.RequestUri,
        Document = document
    }
};
```

A second `[ChunkDocument]`, or a marked property that is not a readable `IngestionDocument`, is reported when the network starts. `[ChunkDocument]` and `[ChunkText]` belong on different types, because one names a document and the other names a string.

## Read the file without a class

With no class to mark, the non-generic `MarkdownSource` writes an `ExpandoObject`. The row has `Document` for the parsed markdown and `StreamMetaData` for the path, the resource count, and the rest of the streaming metadata. `Document` is already the property [Chunk](../chunk/) reads when it holds an `IngestionDocument`, so this row links straight through to the split below.

```csharp
var source = new MarkdownSource("guide.md");
var dest = new MemoryDestination();

source.LinkTo(dest);
Network.Execute(source);

foreach (dynamic row in dest.Data) {
    StreamMetaData metadata = row.StreamMetaData;
    IngestionDocument parsed = row.Document;
    Console.WriteLine(metadata.RequestUri);
    Console.WriteLine(parsed.Identifier);
    Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}
```

`ResultSelector` replaces that pair when the dynamic row should look different, in the same way it builds a class above.

## Read a folder, a URL, or a blob

Each of those row shapes is one element per resource. `Uri` names one resource. `Folder` names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. `RequestCount` follows the sequence.

```csharp
var source = new MarkdownSource {
    Folder = "guides"
};
```

`ResourceType` selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. `GetNextUri` and `HasNextUri` walk a list of addresses when the resources are neither a single path nor a folder. `SkipRows` drops leading lines, and the remainder of that resource is parsed as the one document.

## Split the markdown by heading

This is the step the headings were kept for. Link the dynamic row, or the `IngestionDocument`, straight to [Chunk](../chunk/). A typed row points that transformation at the document with `[ChunkDocument]` or `DocumentSelector`, as set up above.

`HeaderChunker` splits on the headings and keeps the heading path. `Context` on each chunk is that path, such as `Guide > Install`. `Text` is the body under those headings. The limits belong on `IngestionChunkerOptions`, because a chunker you assign is configured through its own options. `MaxTokensPerChunk` and `OverlapTokens` on the transformation apply while `Chunker` is empty.

```csharp
var source = new MarkdownSource<IngestionDocument>("guide.md");
var chunk = new ChunkTransformation<IngestionDocument> {
    Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
        MaxTokensPerChunk = 512,
        OverlapTokens = 50
    })
};

source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);
```

`HeaderChunker`, `IngestionChunkerOptions`, and `IngestionDocument` are in `Microsoft.Extensions.DataIngestion`. `TiktokenTokenizer` is in `Microsoft.ML.Tokenizers`.

On the class from above, `[ChunkDocument]` already names the property, so `DocumentSelector` stays empty:

```csharp
public class Guide {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new MarkdownSource<Guide>("guide.md") {
    ResultSelector = (document, metadata) => new Guide {
        Source = metadata.RequestUri,
        Document = document
    }
};
var chunk = new ChunkTransformation<Guide> {
    Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
        MaxTokensPerChunk = 512,
        OverlapTokens = 50
    })
};
```

## When reading a resource fails

A custom type without `ResultSelector` fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or parsing one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:

```csharp
source.LinkErrorTo(errorDest);
```

Further examples are in the [markdown recipes](/recipes/ai/markdown/).
