# PDF<no value>

{{< callout context="note" icon="outline/info-circle" >}}
One resource leaves as one document, even when a page has no extractable text, and a folder is one document per file in that folder. Plain text, one row per file, remains [Text Document](../text-document/). Markdown remains [Markdown](../markdown/). HTML that is converted to markdown first is [HTML](../html/).
{{< /callout >}}

## One file becomes one document

- **Component**: Source
- **Execution**: One resource in, one document out
- **Package**: `ETLBox.AI.Pdf`

[Markdown](../markdown/) and [HTML](../html/) keep headings. A PDF keeps pages. The file is opened like the other streaming connectors: `Uri`, `Folder`, `ResourceType`, and `RowModificationAction` work the same way, and files, HTTP, and Azure Blob Storage are described in [Shared Features](../streaming-connectors/shared/). The stream is then read to the end and parsed with [PdfPig](https://github.com/UglyToad/PdfPig). Each page becomes a section. Each text block on that page becomes a paragraph, in reading order. Blank blocks are left out. The document `Identifier` is the resource URI, and `PageNumber` is set on the section and on each paragraph. That page structure is what a later section chunker keeps together.

`PdfSource<TOutput>` can write your type, the `IngestionDocument` itself, or an `ExpandoObject`. Install `ETLBox.AI.Pdf`. That package reads the PDF and brings in `ETLBox.AI`, which holds [chunking](../chunk/), chat, and embeddings. The next sections are those three shapes. After that, the same reading covers a folder, a URL, or a blob, and the last step splits the document along the pages.

## Use the parsed document as the row

The simplest row is the document. `PdfSource<IngestionDocument>` emits the parse directly. `Identifier` is the path, the URL, or the blob name. `Sections` are the pages. `Elements` on a section are the paragraphs of that page.

```csharp
var source = new PdfSource<IngestionDocument>("guide.pdf");
var dest = new MemoryDestination<IngestionDocument>();

source.LinkTo(dest);
Network.Execute(source);

foreach (IngestionDocument document in dest.Data) {
    Console.WriteLine(document.Identifier);
    Console.WriteLine(document.Sections.Count);
    Console.WriteLine(document.Sections[0].PageNumber);
    Console.WriteLine(document.Sections[0].Elements[0].Text);
}
```

[Chunk](../chunk/) treats an `IngestionDocument` row as the document, so `DocumentSelector` stays empty. The page split further down uses this form.

## Map the document onto your type

When the rest of the flow needs more than the document, put the document on a class of your own and add the fields you still want, such as the URI. `ResultSelector` receives the parsed document and the `StreamMetaData` and returns the row. Properties you leave unset stay at their default.

Mark the document property with `[ChunkDocument]` when the chunk step below should find it. The attribute belongs on one readable `IngestionDocument`. The source does not read the attribute and does not fill the property. `ResultSelector` still builds the row, and [Chunk](../chunk/) reads the marked property when the network starts. `DocumentSelector` then stays empty. A `DocumentSelector` set on the chunk component is the document that is split. The attribute is what remains when you did not set a selector.

```csharp
public class Article {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new PdfSource<Article>("guide.pdf") {
    ResultSelector = (document, metadata) => new Article {
        Source = metadata.RequestUri,
        Document = document
    }
};
```

A second `[ChunkDocument]`, or a marked property that is not a readable `IngestionDocument`, is reported when the network starts. `[ChunkDocument]` and `[ChunkText]` belong on different types, because one names a document and the other names a string.

## Read the file without a class

With no class to mark, the non-generic `PdfSource` writes an `ExpandoObject`. The row has `Document` for the parsed PDF and `StreamMetaData` for the path, the resource count, and the rest of the streaming metadata. `Document` is already the property [Chunk](../chunk/) reads when it holds an `IngestionDocument`, so this row links straight through to the split below.

```csharp
var source = new PdfSource("guide.pdf");
var dest = new MemoryDestination();

source.LinkTo(dest);
Network.Execute(source);

foreach (dynamic row in dest.Data) {
    StreamMetaData metadata = row.StreamMetaData;
    IngestionDocument parsed = row.Document;
    Console.WriteLine(metadata.RequestUri);
    Console.WriteLine(parsed.Identifier);
    Console.WriteLine(parsed.Sections.Count);
    Console.WriteLine(parsed.Sections[0].Elements[0].Text);
}
```

`ResultSelector` replaces that pair when the dynamic row should look different, in the same way it builds a class above.

## Read a folder, a URL, or a blob

Each of those row shapes is one element per resource. `Uri` names one resource. `Folder` names every file in that folder, and each file becomes its own document. Subfolders stay out of the read. `RequestCount` follows the sequence.

```csharp
var source = new PdfSource {
    Folder = "manuals"
};
```

`ResourceType` selects a file, an HTTP response, or an Azure blob, as on the other streaming sources. One response is one document. `GetNextUri` and `HasNextUri` walk a list of addresses when the resources are neither a single path nor a folder.

`SkipRows` has no lines to drop in a PDF. An assigned value is discarded, the property stays at 0, and the whole file remains the one document.

## Split the PDF into chunks

The pages were kept so this step can follow them. Link the dynamic row, or the `IngestionDocument`, straight to [Chunk](../chunk/). A typed row points that transformation at the document with `[ChunkDocument]` or `DocumentSelector`, as set up above.

Leave `Chunker` empty and the default token chunker cuts the document when a piece reaches `MaxTokensPerChunk`. When a piece should stay inside one page, use `SectionChunker`. Each page is already a section, so that chunker keeps a page together. The limits belong on `IngestionChunkerOptions`, because a chunker you assign is configured through its own options. `MaxTokensPerChunk` and `OverlapTokens` on the transformation apply while `Chunker` is empty.

```csharp
var source = new PdfSource<IngestionDocument>("guide.pdf");
var chunk = new ChunkTransformation<IngestionDocument> {
    Chunker = new SectionChunker(new IngestionChunkerOptions(TiktokenTokenizer.CreateForModel("gpt-4o")) {
        MaxTokensPerChunk = 512,
        OverlapTokens = 0
    })
};

source.LinkTo(chunk).LinkTo(new MemoryDestination<Chunk<IngestionDocument>>());
Network.Execute(source);
```

`SectionChunker`, `IngestionChunkerOptions`, and `IngestionDocument` are in `Microsoft.Extensions.DataIngestion`. `TiktokenTokenizer` is in `Microsoft.ML.Tokenizers`.

On the class from above, `[ChunkDocument]` already names the property, so `DocumentSelector` stays empty. With `Chunker` left empty, the token limits on the transformation apply:

```csharp
public class Guide {
    public string Source { get; set; }

    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

var source = new PdfSource<Guide>("guide.pdf") {
    ResultSelector = (document, metadata) => new Guide {
        Source = metadata.RequestUri,
        Document = document
    }
};
var chunk = new ChunkTransformation<Guide> {
    MaxTokensPerChunk = 100,
    OverlapTokens = 0
};
```

## When reading a resource fails

A custom type without `ResultSelector` fails when the row is built. The source has the document, and the selector is how that type gets created. A failure while reading or parsing one resource follows the normal ETLBox error path. Link an error destination when that resource should be kept aside and the flow should continue:

```csharp
source.LinkErrorTo(errorDest);
```

Further examples are in the [PDF recipes](/recipes/ai/pdf/).
