# HTML<no value>

{{< callout context="note" icon="outline/info-circle" >}}
One resource leaves as one document, even when the HTML is empty, and a folder is one document per file in that folder. Markdown that is already markdown remains [Markdown](../markdown/). Plain text, one row per file, remains [Text Document](../text-document/).
{{< /callout >}}

## One file becomes one document

- **Component**: Source
- **Execution**: One resource in, one document out
- **Package**: `ETLBox.AI.Html`

An HTML page ends up as the same kind of document [Markdown](../markdown/) produces from a `.md` file. The resource is read to the end, converted from HTML to markdown, and then parsed with `MarkdownReader`. Headings in the page stay headings in the document, which is what a later header chunker follows. The document `Identifier` is the resource URI. An empty resource is still one row.

`Uri`, `Folder`, `ResourceType`, encoding, and `RowModificationAction` work as they do on the other streaming connectors. Files, HTTP, and Azure Blob Storage are described in [Shared Features](../streaming-connectors/shared/). `HtmlSource<TOutput>` can write your type, the `IngestionDocument` itself, or an `ExpandoObject`.

Install `ETLBox.AI.Html`. That package reads the HTML and brings in `ETLBox.AI`, which holds [chunking](../chunk/), chat, and embeddings. The conversion itself is the next section. After that, the document leaves as the document, on a class, or on a dynamic row, and the last step splits it on the headings.

## How the HTML is converted

The conversion uses [ReverseMarkdown](https://github.com/mysticmind/reversemarkdown-net), and the markdown that comes out is what `MarkdownReader` parses. `ConverterConfig` is the `ReverseMarkdown.Config` for that first step. A new source removes `script` and `style` elements before the conversion, so a script stays out of the text and the headings are what remain.

Edit that object when you want further ReverseMarkdown options and still want those elements removed. A config you assign yourself starts from ReverseMarkdown's defaults. Put the script and style removal on that object again:

```csharp
var config = new ReverseMarkdown.Config();
config.Preprocess.RemoveScripts().RemoveStyles();

var source = new HtmlSource("guide.html") {
    ConverterConfig = config
};
```

## Use the parsed document as the row

From here the row looks like the one from [Markdown](../markdown/). The simplest form is the document itself. `HtmlSource<IngestionDocument>` emits the parse. `Identifier` is the path, the URL, or the blob name. `Sections` hold the elements the reader produced, including the headings.

```csharp
var source = new HtmlSource<IngestionDocument>("guide.html");
var dest = new MemoryDestination<IngestionDocument>();

source.LinkTo(dest);
Network.Execute(source);

foreach (IngestionDocument document in dest.Data) {
    Console.WriteLine(document.Identifier);
    Console.WriteLine(document.Sections[0].Elements[0].Text);
}
```

[Chunk](../chunk/) treats an `IngestionDocument` row as the document, so `DocumentSelector` stays empty. The heading split further down uses this form.

## Map the document onto your type

When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. `ResultSelector` receives the parsed document and the `StreamMetaData` and returns the row. Properties you leave unset stay at their default.

Mark the document property with `[ChunkDocument]` when the chunk step below should find it. The attribute belongs on one readable `IngestionDocument`. The source does not read the attribute and does not fill the property. `ResultSelector` still builds the row, and [Chunk](../chunk/) reads the marked property when the network starts. `DocumentSelector` then stays empty. A `DocumentSelector` set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.

```csharp
public class Article {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new HtmlSource<Article>("guide.html") {
    ResultSelector = (document, metadata) => new Article {
        Source = metadata.RequestUri,
        Document = document
    }
};
```

A second `[ChunkDocument]`, or a marked property that is not a readable `IngestionDocument`, is reported when the network starts. `[ChunkDocument]` and `[ChunkText]` belong on different types, because one names a document and the other names a string.

## Read the file without a class

With no class to mark, the non-generic `HtmlSource` writes an `ExpandoObject`. The row has `Document` for the parsed HTML and `StreamMetaData` for the path, the resource count, and the rest of the streaming metadata. `Document` is already the property [Chunk](../chunk/) reads when it holds an `IngestionDocument`, so this row links straight through to the split below.

```csharp
var source = new HtmlSource("guide.html");
var dest = new MemoryDestination();

source.LinkTo(dest);
Network.Execute(source);

foreach (dynamic row in dest.Data) {
    StreamMetaData metadata = row.StreamMetaData;
    IngestionDocument parsed = row.Document;
    Console.WriteLine(metadata.RequestUri);
    Console.WriteLine(parsed.Identifier);
    Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}
```

`ResultSelector` replaces that pair when the dynamic row should look different, in the same way it builds a class above.

## Read a folder, a URL, or a blob

Each of those row shapes is one element per resource. `Uri` names one resource. `Folder` names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. `RequestCount` follows the sequence.

```csharp
var source = new HtmlSource {
    Folder = "pages"
};
```

`ResourceType` selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. `GetNextUri` and `HasNextUri` walk a list of addresses when the resources are neither a single path nor a folder. `SkipRows` drops leading lines of the HTML, and the remainder of that resource is converted and parsed as the one document.

## Split the HTML by heading

This is the step the headings were kept for. After the conversion they are markdown headings, so the split is the same one [Markdown](../markdown/) uses. Link the dynamic row, or the `IngestionDocument`, straight to [Chunk](../chunk/). A typed row points that transformation at the document with `[ChunkDocument]` or `DocumentSelector`, as set up above.

`HeaderChunker` splits on those headings and keeps the heading path. `Context` on each chunk is that path, such as `Guide > Install`. `Text` is the body under those headings. The limits belong on `IngestionChunkerOptions`, because a chunker you assign is configured through its own options. `MaxTokensPerChunk` and `OverlapTokens` on the transformation apply while `Chunker` is empty.

```csharp
var source = new HtmlSource<IngestionDocument>("guide.html");
var chunk = new ChunkTransformation<IngestionDocument> {
    Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
        MaxTokensPerChunk = 512,
        OverlapTokens = 50
    })
};

source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);
```

`HeaderChunker`, `IngestionChunkerOptions`, and `IngestionDocument` are in `Microsoft.Extensions.DataIngestion`. `TiktokenTokenizer` is in `Microsoft.ML.Tokenizers`.

On the class from above, `[ChunkDocument]` already names the property, so `DocumentSelector` stays empty:

```csharp
public class Guide {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new HtmlSource<Guide>("guide.html") {
    ResultSelector = (document, metadata) => new Guide {
        Source = metadata.RequestUri,
        Document = document
    }
};
var chunk = new ChunkTransformation<Guide> {
    Chunker = new HeaderChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
        MaxTokensPerChunk = 512,
        OverlapTokens = 50
    })
};
```

## When reading a resource fails

A custom type without `ResultSelector` fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or converting one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:

```csharp
source.LinkErrorTo(errorDest);
```

Further examples are in the [HTML recipes](/recipes/ai/html/).
