# PDF Source<no value>

`PdfSource` reads one PDF resource as one `IngestionDocument`. A file, a web response, or a blob is one element. A folder is one element per file. Each page becomes a section, and each text block on that page becomes a paragraph. The non-generic source writes `Document` and `StreamMetaData` on an `ExpandoObject`.

The examples below read `guide.pdf`. The file has two pages.

Page 1:

```
Install ETLBox
Run dotnet add package ETLBox.AI.
```

Page 2:

```
Other
Unrelated page body.
```

## Read one file

The whole file is one dynamic row. `Document` is the parsed PDF. `StreamMetaData` carries the path and the resource count. `Sections` are the pages.

```C#
string sourceFile = "guide.pdf";

var source = new PdfSource(sourceFile);
var dest = new MemoryDestination();

source.LinkTo(dest);
Network.Execute(source);

foreach (dynamic row in dest.Data) {
    StreamMetaData metadata = row.StreamMetaData;
    IngestionDocument parsed = row.Document;
    Console.WriteLine($"RequestUri:{metadata.RequestUri}");
    Console.WriteLine($"Identifier:{parsed.Identifier}");
    Console.WriteLine($"Pages:{parsed.Sections.Count}");
    Console.WriteLine($"Page 1:{parsed.Sections[0].Elements[0].Text}");
}

//Outputs
//RequestUri:guide.pdf
//Identifier:guide.pdf
//Pages:2
//Page 1:Install ETLBox
//Run dotnet add package ETLBox.AI.
```

## Build the row yourself

`ResultSelector` gets the parsed document and the `StreamMetaData`. Fields such as `Source` are set here.

```C#
public class Article
{
    public string Source { get; set; }
    public IngestionDocument Document { get; set; }
}

string sourceFile = "guide.pdf";

var source = new PdfSource<Article>(sourceFile) {
    ResultSelector = (document, metadata) => new Article {
        Source = metadata.RequestUri,
        Document = document
    }
};
var dest = new MemoryDestination<Article>();

source.LinkTo(dest);
Network.Execute(source);

foreach (Article row in dest.Data)
    Console.WriteLine($"Source:{row.Source} Pages:{row.Document.Sections.Count}");

//Outputs
//Source:guide.pdf Pages:2
```

## Read the document itself

`PdfSource<IngestionDocument>` emits the parsed document. The row has no separate property for the path, because `Identifier` already is that URI. `PageNumber` on each section is the page.

```C#
string sourceFile = "guide.pdf";

var source = new PdfSource<IngestionDocument>(sourceFile);
var dest = new MemoryDestination<IngestionDocument>();

source.LinkTo(dest);
Network.Execute(source);

foreach (IngestionDocument document in dest.Data)
    Console.WriteLine($"Identifier:{document.Identifier} Pages:{document.Sections.Count}");

//Outputs
//Identifier:guide.pdf Pages:2
```

## Split your own type

`[ChunkDocument]` names the document property. `ResultSelector` still builds the row, because the source does not fill a custom type on its own. With `Chunker` left empty, `MaxTokensPerChunk` and `OverlapTokens` cut the text.

```C#
public class Guide
{
    public string Source { get; set; }
    [ChunkDocument]
    public IngestionDocument Document { get; set; }
}

string sourceFile = "guide.pdf";

var source = new PdfSource<Guide>(sourceFile) {
    ResultSelector = (document, metadata) => new Guide {
        Source = metadata.RequestUri,
        Document = document
    }
};
//No DocumentSelector needed - the attribute defines the document property
var chunk = new ChunkTransformation<Guide> {
    MaxTokensPerChunk = 100,
    OverlapTokens = 0
};
var dest = new MemoryDestination<Chunk<Guide>>();

source.LinkTo(chunk);
chunk.LinkTo(dest);
Network.Execute(source);

foreach (Chunk<Guide> row in dest.Data)
    Console.WriteLine($"Source:{row.Source.Source} Text:{row.Text}");

//Outputs
//Source:guide.pdf Text:Install ETLBox Run dotnet add package ETLBox.AI.
//Source:guide.pdf Text:Other Unrelated page body.
```
