Files
docling-eval/tests/test_predictions_visualizer.py
373f959633 feat: Visualizer tool and command for datasets (#186)
* chore: Move the teds.py inside the subdir evaluators/table

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce the external_predictions_path in BaseEvaluator and dummy entries in all evaluators.
Extend the CLI to support the --external-predictions-path

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend test_dataset_builder.py to save document predictions in various formats

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend MarkDownTextEvaluator to support external_predictions_path. Add unit test

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend LayoutEvaluator to support external_predictions_path. Add unit test.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Add missing pytest dependencies in tests

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Fix loading the external predictions in LayoutEvaluator

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce external predictions in DocStructureEvaluator. Add unit test.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the TableEvaluator to support external predictions. Add unit test

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the KeyValueEvaluator to support external predictions. Add unit test.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the PixelLayoutEvaluator to support external predictions. Add unit test

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the BboxTextEvaluator to support external predictions. Add unit test

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Disable the OCREvaluator when using the external predictions

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Fixing guard for external predictions in TimingsEvaluator, ReadingOrderEvaluator. Fix main

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Export the doctag files with the correct file extension

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Refactor the ExternalDoclingDocumentLoader to properly load a DoclingDocument from doctags and
the GT image.
- Introduce the staticmethod load_doctags() which covers all cases on page image loading.
- Refactor the FilePredictionProvider to use the load_doctags() from ExternalDoclingDocumentLoader.
- Refactor all evaluators to use the new ExternalDoclingDocumentLoader.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Rename code file as external_docling_document_loader.py

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Fix typo

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce examples how to evaluate using external predictions using the API and the CLI.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Prediction vizualizer

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* Update docling_eval/utils/external_predictions_visualizer.py

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Signed-off-by: Christoph Auer <60343111+cau-git@users.noreply.github.com>

* feat: Update examples bash script to demonstrate visualisations on external predictions

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

---------

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Signed-off-by: Christoph Auer <60343111+cau-git@users.noreply.github.com>
Co-authored-by: Nikos Livathinos <nli@zurich.ibm.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-12-09 14:47:43 +01:00

65 lines
2.1 KiB
Python

from pathlib import Path
import pytest
from datasets import load_dataset
from docling_eval.datamodels.dataset_record import DatasetRecordWithPrediction
from docling_eval.utils.external_predictions_visualizer import PredictionsVisualizer
def _first_doc_id(parquet_root: Path) -> str:
split_files = sorted((parquet_root / "test").glob("*.parquet"))
ds = load_dataset(
"parquet", data_files={"test": [str(path) for path in split_files]}
)
record = DatasetRecordWithPrediction.model_validate(ds["test"][0])
return record.doc_id
@pytest.mark.dependency(
depends=["tests/test_dataset_builder.py::test_run_dpbench_e2e"],
scope="session",
)
def test_predictions_visualizer_with_embedded_predictions() -> None:
dataset_dir = Path("scratch/DPBench/eval_dataset_e2e")
output_dir = Path("scratch/DPBench/visualizer_tests/embedded")
output_dir.mkdir(parents=True, exist_ok=True)
visualizer = PredictionsVisualizer(visualizations_dir=output_dir)
visualizer.create_visualizations(
dataset_dir=dataset_dir,
split="test",
begin_index=0,
end_index=1,
)
doc_id = _first_doc_id(dataset_dir)
layout_file = output_dir / f"{doc_id}_layout.html"
assert layout_file.is_file()
@pytest.mark.dependency(
depends=["tests/test_dataset_builder.py::test_run_dpbench_e2e"],
scope="session",
)
def test_predictions_visualizer_with_external_predictions() -> None:
gt_dir = Path("scratch/DPBench/gt_dataset")
external_predictions_dir = Path("scratch/DPBench/predicted_documents/json")
output_dir = Path("scratch/DPBench/visualizer_tests/external")
output_dir.mkdir(parents=True, exist_ok=True)
visualizer = PredictionsVisualizer(
visualizations_dir=output_dir,
external_predictions_dir=external_predictions_dir,
)
visualizer.create_visualizations(
dataset_dir=gt_dir,
split="test",
begin_index=0,
end_index=1,
)
doc_id = _first_doc_id(gt_dir)
layout_file = output_dir / f"{doc_id}_layout.html"
assert layout_file.is_file()