docling-ibm-models

Name	docling-ibm-models JSON
Version	3.8.1 JSON
	download
home_page	None
Summary	This package contains the AI models used by the Docling PDF conversion package
upload_time	2025-07-10 12:45:29
maintainer	None
docs_url	None
author	None
requires_python	<4.0,>=3.9
license	None
keywords	docling convert document pdf layout model segmentation table structure table former
VCS
bugtrack_url
requirements	No requirements were recorded.
Travis-CI	No Travis.
coveralls test coverage	No coveralls.

            [![PyPI version](https://img.shields.io/pypi/v/docling-ibm-models)](https://pypi.org/project/docling-ibm-models/)
[![PyPI - Python Version](https://img.shields.io/pypi/pyversions/docling-ibm-models)](https://pypi.org/project/docling-ibm-models/)
[![uv](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json)](https://github.com/astral-sh/uv)
[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)
[![Imports: isort](https://img.shields.io/badge/%20imports-isort-%231674b1?style=flat&labelColor=ef8336)](https://pycqa.github.io/isort/)
[![pre-commit](https://img.shields.io/badge/pre--commit-enabled-brightgreen?logo=pre-commit&logoColor=white)](https://github.com/pre-commit/pre-commit)
[![Models on Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-blue)](https://huggingface.co/ds4sd/docling-models/)
[![License MIT](https://img.shields.io/github/license/ds4sd/deepsearch-toolkit)](https://opensource.org/licenses/MIT)

# Docling IBM models

AI modules to support the Docling PDF document conversion project.

- TableFormer is an AI module that recognizes the structure of a table and the bounding boxes of the table content.
- Layout model is an AI model that provides among other things ability to detect tables on the page. This package contains inference code for Layout model.


## Pipeline Overview
![Architecture](docs/tablemodel_overview_color.png)

## Datasets
Below we list datasets used with their description, source, and ***"TableFormer Format"***. The TableFormer Format is our processed version of the version of the original format to work with the dataloader out of the box, and to augment the dataset when necassary to add missing groundtruth (bounding boxes for empty cells).


| Name        | Description      | URL |
| ------------- |:-------------:|----|
| PubTabNet | PubTabNet contains heterogeneous tables in both image and HTML format, 516k+ tables in the PubMed Central Open Access Subset  | [PubTabNet](https://developer.ibm.com/exchanges/data/all/pubtabnet/) |
| FinTabNet| A dataset for Financial Report Tables with corresponding ground truth location and structure. 112k+ tables included.| [FinTabNet](https://developer.ibm.com/exchanges/data/all/fintabnet/) |
| TableBank| TableBank is a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet, contains 417K high-quality labeled tables. | [TableBank](https://github.com/doc-analysis/TableBank) |

## Models

### TableModel04:
![TableModel04](docs/tbm04.png)
**TableModel04rs (OTSL)** is our SOTA method that using transformers in order to predict table structure and bounding box.


## Configuration file

Example configuration can be found inside test `tests/test_tf_predictor.py`
These are the main sections of the configuration file:

- `dataset`: The directory for prepared data and the parameters used during the data loading.
- `model`: The type, name and hyperparameters of the model. Also the directory to save/load the
  trained checkpoint files.
- `train`: Parameters for the training of the model.
- `predict`: Parameters for the evaluation of the model.
- `dataset_wordmap`: Very important part that contains token maps.


## Model weights

You can download the model weights and config files from the links:

- [TableFormer Checkpoint](https://huggingface.co/ds4sd/docling-models/tree/main/model_artifacts/tableformer)
- [beehive_v0.0.5](https://huggingface.co/ds4sd/docling-models/tree/main/model_artifacts/layout/beehive_v0.0.5)


## Inference Tests

You can run the inference tests for the models with:

```
python -m pytest tests/
```

This will also generate prediction and matching visualizations that can be found here:
`tests\test_data\viz\`

Visualization outlines:
- `Light Pink`: border of recognized table
- `Grey`: OCR cells
- `Green`: prediction bboxes
- `Red`: OCR cells matched with prediction
- `Blue`: Post processed, match
- `Bold Blue`: column header
- `Bold Magenta`: row header
- `Bold Brown`: section row (if table have one)


## Demo

A demo application allows to apply the `LayoutPredictor` on a directory `<input_dir>` that contains
`png` images and visualize the predictions inside another directory `<viz_dir>`.

First download the model weights (see above), then run:
```
python -m demo.demo_layout_predictor -i <input_dir> -v <viz_dir>
```

e.g.
```
python -m demo.demo_layout_predictor -i tests/test_data/samples -v viz/
```

Raw data

            {
    "_id": null,
    "home_page": null,
    "name": "docling-ibm-models",
    "maintainer": null,
    "docs_url": null,
    "requires_python": "<4.0,>=3.9",
    "maintainer_email": null,
    "keywords": "docling, convert, document, pdf, layout model, segmentation, table structure, table former",
    "author": null,
    "author_email": "Nikos Livathinos <nli@zurich.ibm.com>, Maxim Lysak <mly@zurich.ibm.com>, Ahmed Nassar <ahn@zurich.ibm.com>, Christoph Auer <cau@zurich.ibm.com>, Michele Dolfi <dol@zurich.ibm.com>, Peter Staar <taa@zurich.ibm.com>",
    "download_url": "https://files.pythonhosted.org/packages/95/9b/67d04fa085f1a99a9ecee5a13d4bd472c82a14c260eedd169feb2c63649b/docling_ibm_models-3.8.1.tar.gz",
    "platform": null,
    "description": "[![PyPI version](https://img.shields.io/pypi/v/docling-ibm-models)](https://pypi.org/project/docling-ibm-models/)\n[![PyPI - Python Version](https://img.shields.io/pypi/pyversions/docling-ibm-models)](https://pypi.org/project/docling-ibm-models/)\n[![uv](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json)](https://github.com/astral-sh/uv)\n[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)\n[![Imports: isort](https://img.shields.io/badge/%20imports-isort-%231674b1?style=flat&labelColor=ef8336)](https://pycqa.github.io/isort/)\n[![pre-commit](https://img.shields.io/badge/pre--commit-enabled-brightgreen?logo=pre-commit&logoColor=white)](https://github.com/pre-commit/pre-commit)\n[![Models on Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-blue)](https://huggingface.co/ds4sd/docling-models/)\n[![License MIT](https://img.shields.io/github/license/ds4sd/deepsearch-toolkit)](https://opensource.org/licenses/MIT)\n\n# Docling IBM models\n\nAI modules to support the Docling PDF document conversion project.\n\n- TableFormer is an AI module that recognizes the structure of a table and the bounding boxes of the table content.\n- Layout model is an AI model that provides among other things ability to detect tables on the page. This package contains inference code for Layout model.\n\n\n## Pipeline Overview\n![Architecture](docs/tablemodel_overview_color.png)\n\n## Datasets\nBelow we list datasets used with their description, source, and ***\"TableFormer Format\"***. The TableFormer Format is our processed version of the version of the original format to work with the dataloader out of the box, and to augment the dataset when necassary to add missing groundtruth (bounding boxes for empty cells).\n\n\n| Name        | Description      | URL |\n| ------------- |:-------------:|----|\n| PubTabNet | PubTabNet contains heterogeneous tables in both image and HTML format, 516k+ tables in the PubMed Central Open Access Subset  | [PubTabNet](https://developer.ibm.com/exchanges/data/all/pubtabnet/) |\n| FinTabNet| A dataset for Financial Report Tables with corresponding ground truth location and structure. 112k+ tables included.| [FinTabNet](https://developer.ibm.com/exchanges/data/all/fintabnet/) |\n| TableBank| TableBank is a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet, contains 417K high-quality labeled tables. | [TableBank](https://github.com/doc-analysis/TableBank) |\n\n## Models\n\n### TableModel04:\n![TableModel04](docs/tbm04.png)\n**TableModel04rs (OTSL)** is our SOTA method that using transformers in order to predict table structure and bounding box.\n\n\n## Configuration file\n\nExample configuration can be found inside test `tests/test_tf_predictor.py`\nThese are the main sections of the configuration file:\n\n- `dataset`: The directory for prepared data and the parameters used during the data loading.\n- `model`: The type, name and hyperparameters of the model. Also the directory to save/load the\n  trained checkpoint files.\n- `train`: Parameters for the training of the model.\n- `predict`: Parameters for the evaluation of the model.\n- `dataset_wordmap`: Very important part that contains token maps.\n\n\n## Model weights\n\nYou can download the model weights and config files from the links:\n\n- [TableFormer Checkpoint](https://huggingface.co/ds4sd/docling-models/tree/main/model_artifacts/tableformer)\n- [beehive_v0.0.5](https://huggingface.co/ds4sd/docling-models/tree/main/model_artifacts/layout/beehive_v0.0.5)\n\n\n## Inference Tests\n\nYou can run the inference tests for the models with:\n\n```\npython -m pytest tests/\n```\n\nThis will also generate prediction and matching visualizations that can be found here:\n`tests\\test_data\\viz\\`\n\nVisualization outlines:\n- `Light Pink`: border of recognized table\n- `Grey`: OCR cells\n- `Green`: prediction bboxes\n- `Red`: OCR cells matched with prediction\n- `Blue`: Post processed, match\n- `Bold Blue`: column header\n- `Bold Magenta`: row header\n- `Bold Brown`: section row (if table have one)\n\n\n## Demo\n\nA demo application allows to apply the `LayoutPredictor` on a directory `<input_dir>` that contains\n`png` images and visualize the predictions inside another directory `<viz_dir>`.\n\nFirst download the model weights (see above), then run:\n```\npython -m demo.demo_layout_predictor -i <input_dir> -v <viz_dir>\n```\n\ne.g.\n```\npython -m demo.demo_layout_predictor -i tests/test_data/samples -v viz/\n```\n",
    "bugtrack_url": null,
    "license": null,
    "summary": "This package contains the AI models used by the Docling PDF conversion package",
    "version": "3.8.1",
    "project_urls": {
        "changelog": "https://github.com/docling-project/docling-ibm-models/blob/main/CHANGELOG.md",
        "homepage": "https://github.com/docling-project/docling-ibm-models",
        "issues": "https://github.com/docling-project/docling-ibm-models/issues",
        "repository": "https://github.com/docling-project/docling-ibm-models"
    },
    "split_keywords": [
        "docling",
        " convert",
        " document",
        " pdf",
        " layout model",
        " segmentation",
        " table structure",
        " table former"
    ],
    "urls": [
        {
            "comment_text": null,
            "digests": {
                "blake2b_256": "983c85f75b70b5ea277e2f07c4c50c8037c9be61dc272ea55c9576716a7358f7",
                "md5": "adfada95005802158b1a5dd01f7c7e3b",
                "sha256": "91ef82c63d497b6b91e8c85e0b7a84f75ec38e3495374bfe6112a71081bf8c2a"
            },
            "downloads": -1,
            "filename": "docling_ibm_models-3.8.1-py3-none-any.whl",
            "has_sig": false,
            "md5_digest": "adfada95005802158b1a5dd01f7c7e3b",
            "packagetype": "bdist_wheel",
            "python_version": "py3",
            "requires_python": "<4.0,>=3.9",
            "size": 86120,
            "upload_time": "2025-07-10T12:45:27",
            "upload_time_iso_8601": "2025-07-10T12:45:27.592786Z",
            "url": "https://files.pythonhosted.org/packages/98/3c/85f75b70b5ea277e2f07c4c50c8037c9be61dc272ea55c9576716a7358f7/docling_ibm_models-3.8.1-py3-none-any.whl",
            "yanked": false,
            "yanked_reason": null
        },
        {
            "comment_text": null,
            "digests": {
                "blake2b_256": "959b67d04fa085f1a99a9ecee5a13d4bd472c82a14c260eedd169feb2c63649b",
                "md5": "114026f7fa06272d2737a185cf4ccf33",
                "sha256": "721982f8b64b8bd9acdfb96cc9e0e9162b51a0f318b6c5a9e814f8d1b5b537aa"
            },
            "downloads": -1,
            "filename": "docling_ibm_models-3.8.1.tar.gz",
            "has_sig": false,
            "md5_digest": "114026f7fa06272d2737a185cf4ccf33",
            "packagetype": "sdist",
            "python_version": "source",
            "requires_python": "<4.0,>=3.9",
            "size": 86083,
            "upload_time": "2025-07-10T12:45:29",
            "upload_time_iso_8601": "2025-07-10T12:45:29.233752Z",
            "url": "https://files.pythonhosted.org/packages/95/9b/67d04fa085f1a99a9ecee5a13d4bd472c82a14c260eedd169feb2c63649b/docling_ibm_models-3.8.1.tar.gz",
            "yanked": false,
            "yanked_reason": null
        }
    ],
    "upload_time": "2025-07-10 12:45:29",
    "github": true,
    "gitlab": false,
    "bitbucket": false,
    "codeberg": false,
    "github_user": "docling-project",
    "github_project": "docling-ibm-models",
    "travis_ci": false,
    "coveralls": false,
    "github_actions": true,
    "lcname": "docling-ibm-models"
}

None