docling/docs/usage.md

## Conversion

### Convert a single document

To convert individual PDF documents, use `convert()`, for example:

```python
from docling.document_converter import DocumentConverter

source = "https://arxiv.org/pdf/2408.09869"  # PDF path or URL
converter = DocumentConverter()
result = converter.convert(source)
print(result.document.export_to_markdown())  # output: "### Docling Technical Report[...]"
```

### CLI

You can also use Docling directly from your command line to convert individual files —be it local or by URL— or whole directories.

A simple example would look like this:
```console
docling https://arxiv.org/pdf/2206.01062
```

To see all available options (export formats etc.) run `docling --help`. More details in the [CLI reference page](./reference/cli.md).

### Advanced options

#### Adjust pipeline features

The example file [custom_convert.py](./examples/custom_convert.py) contains multiple ways
one can adjust the conversion pipeline and features.


##### Control PDF table extraction options

You can control if table structure recognition should map the recognized structure back to PDF cells (default) or use text cells from the structure prediction itself.
This can improve output quality if you find that multiple columns in extracted tables are erroneously merged into one.


```python
from docling.datamodel.base_models import InputFormat
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions

pipeline_options = PdfPipelineOptions(do_table_structure=True)
pipeline_options.table_structure_options.do_cell_matching = False  # uses text cells predicted from table structure model

doc_converter = DocumentConverter(
    format_options={
        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
    }
)
```

Since docling 1.16.0: You can control which TableFormer mode you want to use. Choose between `TableFormerMode.FAST` (default) and `TableFormerMode.ACCURATE` (better, but slower) to receive better quality with difficult table structures.

```python
from docling.datamodel.base_models import InputFormat
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions, TableFormerMode

pipeline_options = PdfPipelineOptions(do_table_structure=True)
pipeline_options.table_structure_options.mode = TableFormerMode.ACCURATE  # use more accurate TableFormer model

doc_converter = DocumentConverter(
    format_options={
        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
    }
)
```

##### Provide specific artifacts path

By default, artifacts such as models are downloaded automatically upon first usage. If you would prefer to use a local path where the artifacts have been explicitly prefetched, you can do that as follows:

```python
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.pipeline.standard_pdf_pipeline import StandardPdfPipeline

# # to explicitly prefetch:
# artifacts_path = StandardPdfPipeline.download_models_hf()

artifacts_path = "/local/path/to/artifacts"

pipeline_options = PdfPipelineOptions(artifacts_path=artifacts_path)
doc_converter = DocumentConverter(
    format_options={
        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
    }
)
```

#### Impose limits on the document size

You can limit the file size and number of pages which should be allowed to process per document:

```python
from pathlib import Path
from docling.document_converter import DocumentConverter

source = "https://arxiv.org/pdf/2408.09869"
converter = DocumentConverter()
result = converter.convert(source, max_num_pages=100, max_file_size=20971520)
```

#### Convert from binary PDF streams

You can convert PDFs from a binary stream instead of from the filesystem as follows:

```python
from io import BytesIO
from docling.datamodel.base_models import DocumentStream
from docling.document_converter import DocumentConverter

buf = BytesIO(your_binary_stream)
source = DocumentStream(name="my_doc.pdf", stream=buf)
converter = DocumentConverter()
result = converter.convert(source)
```

#### Limit resource usage

You can limit the CPU threads used by Docling by setting the environment variable `OMP_NUM_THREADS` accordingly. The default setting is using 4 CPU threads.


#### Use specific backend converters

!!! note

    This section discusses directly invoking a [backend](./concepts/architecture.md),
    i.e. using a low-level API. This should only be done when necessary. For most cases,
    using a `DocumentConverter` (high-level API) as discussed in the sections above
    should suffice — and is the recommended way.

By default, Docling will try to identify the document format to apply the appropriate conversion backend (see the list of [supported formats](./supported_formats.md)).
You can restrict the `DocumentConverter` to a set of allowed document formats, as shown in the [Multi-format conversion](./examples/run_with_formats.py) example.
Alternatively, you can also use the specific backend that matches your document content. For instance, you can use `HTMLDocumentBackend` for HTML pages:

```python
import urllib.request
from io import BytesIO
from docling.backend.html_backend import HTMLDocumentBackend
from docling.datamodel.base_models import InputFormat
from docling.datamodel.document import InputDocument

url = "https://en.wikipedia.org/wiki/Duck"
text = urllib.request.urlopen(url).read()
in_doc = InputDocument(
    path_or_stream=BytesIO(text),
    format=InputFormat.HTML,
    backend=HTMLDocumentBackend,
    filename="duck.html",
)
backend = HTMLDocumentBackend(in_doc=in_doc, path_or_stream=BytesIO(text))
dl_doc = backend.convert()
print(dl_doc.export_to_markdown())
```

## Chunking

You can chunk a Docling document using a [chunker](concepts/chunking.md), such as a
`HybridChunker`, as shown below (for more details check out
[this example](examples/hybrid_chunking.ipynb)):

```python
from docling.document_converter import DocumentConverter
from docling.chunking import HybridChunker

conv_res = DocumentConverter().convert("https://arxiv.org/pdf/2206.01062")
doc = conv_res.document

chunker = HybridChunker(tokenizer="BAAI/bge-small-en-v1.5")  # set tokenizer as needed
chunk_iter = chunker.chunk(doc)
```

An example chunk would look like this:

```python
print(list(chunk_iter)[11])
# {
#   "text": "In this paper, we present the DocLayNet dataset. [...]",
#   "meta": {
#     "doc_items": [{
#       "self_ref": "#/texts/28",
#       "label": "text",
#       "prov": [{
#         "page_no": 2,
#         "bbox": {"l": 53.29, "t": 287.14, "r": 295.56, "b": 212.37, ...},
#       }], ...,
#     }, ...],
#     "headings": ["1 INTRODUCTION"],
#   }
# }
```
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								## Conversion
 								### Convert a single document
-												docs: correct spelling of 'individual' (#219)

Signed-off-by: Vicky Sekhon <114193273+VickySekhon@users.noreply.github.com>
											
										
										
											2024-11-04 08:27:02 -05:00
+								To convert individual PDF documents, use `convert()`, for example:
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
 								```python
 								from docling.document_converter import DocumentConverter
 								source = "https://arxiv.org/pdf/2408.09869"  # PDF path or URL
 								converter = DocumentConverter()
 								result = converter.convert(source)
 								print(result.document.export_to_markdown())  # output: "### Docling Technical Report[...]"
 								```
 								### CLI
 								You can also use Docling directly from your command line to convert individual files —be it local or by URL— or whole directories.
 								A simple example would look like this:
 								```console
 								docling https://arxiv.org/pdf/2206.01062
 								```
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								To see all available options (export formats etc.) run `docling --help`. More details in the [CLI reference page](./reference/cli.md).
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
 								### Advanced options
 								#### Adjust pipeline features
 								The example file [custom_convert.py](./examples/custom_convert.py) contains multiple ways
 								one can adjust the conversion pipeline and features.
 								##### Control PDF table extraction options
 								You can control if table structure recognition should map the recognized structure back to PDF cells (default) or use text cells from the structure prediction itself.
 								This can improve output quality if you find that multiple columns in extracted tables are erroneously merged into one.
 								```python
 								from docling.datamodel.base_models import InputFormat
 								from docling.document_converter import DocumentConverter, PdfFormatOption
 								from docling.datamodel.pipeline_options import PdfPipelineOptions
 								pipeline_options = PdfPipelineOptions(do_table_structure=True)
 								pipeline_options.table_structure_options.do_cell_matching = False  # uses text cells predicted from table structure model
 								doc_converter = DocumentConverter(
 								    format_options={
 								        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
 								    }
 								)
 								```
 								Since docling 1.16.0: You can control which TableFormer mode you want to use. Choose between `TableFormerMode.FAST` (default) and `TableFormerMode.ACCURATE` (better, but slower) to receive better quality with difficult table structures.
 								```python
 								from docling.datamodel.base_models import InputFormat
 								from docling.document_converter import DocumentConverter, PdfFormatOption
 								from docling.datamodel.pipeline_options import PdfPipelineOptions, TableFormerMode
 								pipeline_options = PdfPipelineOptions(do_table_structure=True)
 								pipeline_options.table_structure_options.mode = TableFormerMode.ACCURATE  # use more accurate TableFormer model
 								doc_converter = DocumentConverter(
 								    format_options={
 								        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
 								    }
 								)
 								```
-												docs: add explicit artifacts path example (#224)

* docs: add explicit artifacts path example

[skip ci]

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

* minor docs fix

[skip ci]

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

* touch to trigger needed checks

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

---------

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-11-04 14:27:56 +01:00
+								##### Provide specific artifacts path
 								By default, artifacts such as models are downloaded automatically upon first usage. If you would prefer to use a local path where the artifacts have been explicitly prefetched, you can do that as follows:
 								```python
 								from docling.datamodel.base_models import InputFormat
 								from docling.datamodel.pipeline_options import PdfPipelineOptions
 								from docling.document_converter import DocumentConverter, PdfFormatOption
 								from docling.pipeline.standard_pdf_pipeline import StandardPdfPipeline
 								# # to explicitly prefetch:
 								# artifacts_path = StandardPdfPipeline.download_models_hf()
 								artifacts_path = "/local/path/to/artifacts"
 								pipeline_options = PdfPipelineOptions(artifacts_path=artifacts_path)
 								doc_converter = DocumentConverter(
 								    format_options={
 								        InputFormat.PDF: PdfFormatOption(pipeline_options=pipeline_options)
 								    }
 								)
 								```
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								#### Impose limits on the document size
 								You can limit the file size and number of pages which should be allowed to process per document:
 								```python
 								from pathlib import Path
 								from docling.document_converter import DocumentConverter
 								source = "https://arxiv.org/pdf/2408.09869"
 								converter = DocumentConverter()
 								result = converter.convert(source, max_num_pages=100, max_file_size=20971520)
 								```
 								#### Convert from binary PDF streams
 								You can convert PDFs from a binary stream instead of from the filesystem as follows:
 								```python
 								from io import BytesIO
 								from docling.datamodel.base_models import DocumentStream
 								from docling.document_converter import DocumentConverter
 								buf = BytesIO(your_binary_stream)
-												docs: fix parameter in usage.md (#332)

Signed-off-by: Carl Senze <carl.senze@aleph-alpha.com>
Co-authored-by: Carl Senze <carl.senze@aleph-alpha.com>
											
										
										
											2024-11-15 09:24:15 +01:00
+								source = DocumentStream(name="my_doc.pdf", stream=buf)
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								converter = DocumentConverter()
 								result = converter.convert(source)
 								```
 								#### Limit resource usage
 								You can limit the CPU threads used by Docling by setting the environment variable `OMP_NUM_THREADS` accordingly. The default setting is using 4 CPU threads.
-												docs: description of supported formats and backends (#788)

* chore: remove type-ignore marks for attaching text to non GroupItems

After commit b74208 of docling-core, text items can be attached to any NodeItem
and therefore the ignore[arg-type] type marks can be removed.

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* test: remove unnecessary imports

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add documentation on supported formats and backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add notebook example with XML backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

---------

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>
											
										
										
											2025-01-26 08:10:33 +01:00
+								#### Use specific backend converters
-												docs: document Docling JSON parsing (#819)

* docs: document Docling JSON parsing

Also:
- factored out and expanded supported formats
- reorged feature list

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

* update feature list, minor fixes

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

---------

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2025-01-28 13:23:30 +01:00
+								!!! note
 								    This section discusses directly invoking a [backend](./concepts/architecture.md),
 								    i.e. using a low-level API. This should only be done when necessary. For most cases,
 								    using a `DocumentConverter` (high-level API) as discussed in the sections above
 								    should suffice — and is the recommended way.
 								By default, Docling will try to identify the document format to apply the appropriate conversion backend (see the list of [supported formats](./supported_formats.md)).
-												docs: description of supported formats and backends (#788)

* chore: remove type-ignore marks for attaching text to non GroupItems

After commit b74208 of docling-core, text items can be attached to any NodeItem
and therefore the ignore[arg-type] type marks can be removed.

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* test: remove unnecessary imports

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add documentation on supported formats and backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add notebook example with XML backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

---------

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>
											
										
										
											2025-01-26 08:10:33 +01:00
+								You can restrict the `DocumentConverter` to a set of allowed document formats, as shown in the [Multi-format conversion](./examples/run_with_formats.py) example.
 								Alternatively, you can also use the specific backend that matches your document content. For instance, you can use `HTMLDocumentBackend` for HTML pages:
 								```python
 								import urllib.request
 								from io import BytesIO
 								from docling.backend.html_backend import HTMLDocumentBackend
 								from docling.datamodel.base_models import InputFormat
 								from docling.datamodel.document import InputDocument
 								url = "https://en.wikipedia.org/wiki/Duck"
 								text = urllib.request.urlopen(url).read()
 								in_doc = InputDocument(
 								    path_or_stream=BytesIO(text),
 								    format=InputFormat.HTML,
 								    backend=HTMLDocumentBackend,
 								    filename="duck.html",
 								)
 								backend = HTMLDocumentBackend(in_doc=in_doc, path_or_stream=BytesIO(text))
-												docs: document Docling JSON parsing (#819)

* docs: document Docling JSON parsing

Also:
- factored out and expanded supported formats
- reorged feature list

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

* update feature list, minor fixes

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>

---------

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2025-01-28 13:23:30 +01:00
+								dl_doc = backend.convert()
 								print(dl_doc.export_to_markdown())
-												docs: description of supported formats and backends (#788)

* chore: remove type-ignore marks for attaching text to non GroupItems

After commit b74208 of docling-core, text items can be attached to any NodeItem
and therefore the ignore[arg-type] type marks can be removed.

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* test: remove unnecessary imports

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add documentation on supported formats and backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

* docs: add notebook example with XML backends

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>

---------

Signed-off-by: Cesar Berrospi Ramis <75900930+ceberam@users.noreply.github.com>
											
										
										
											2025-01-26 08:10:33 +01:00
+								```
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								## Chunking
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								You can chunk a Docling document using a [chunker](concepts/chunking.md), such as a
 								`HybridChunker`, as shown below (for more details check out
 								[this example](examples/hybrid_chunking.ipynb)):
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
 								```python
 								from docling.document_converter import DocumentConverter
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								from docling.chunking import HybridChunker
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
 								conv_res = DocumentConverter().convert("https://arxiv.org/pdf/2206.01062")
 								doc = conv_res.document
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								chunker = HybridChunker(tokenizer="BAAI/bge-small-en-v1.5")  # set tokenizer as needed
 								chunk_iter = chunker.chunk(doc)
 								```
 								An example chunk would look like this:
 								```python
 								print(list(chunk_iter)[11])
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								# {
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								#   "text": "In this paper, we present the DocLayNet dataset. [...]",
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								#   "meta": {
 								#     "doc_items": [{
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								#       "self_ref": "#/texts/28",
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								#       "label": "text",
 								#       "prov": [{
 								#         "page_no": 2,
-												docs: update chunking usage docs, minor reorg (#550)

Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-12-10 16:03:02 +01:00
+								#         "bbox": {"l": 53.29, "t": 287.14, "r": 295.56, "b": 212.37, ...},
 								#       }], ...,
 								#     }, ...],
 								#     "headings": ["1 INTRODUCTION"],
-												docs: add use docling (#150)


---------

Signed-off-by: Michele Dolfi <dol@zurich.ibm.com>
Signed-off-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
Co-authored-by: Panos Vagenas <35837085+vagenas@users.noreply.github.com>
											
										
										
											2024-10-17 18:14:48 +02:00
+								#   }
 								# }
 								```