Commit Graph
2 Commits
Author SHA1 Message Date
Nikos Livathinos 04fe2d916f feat: Improvements for the MultiEvaluator (#95)
* fix: MultiEvaluator fix minor logging issue

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Improve code comments

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Refactor the MultiEvaluator to allow arbitrary experiment names for the benchmark subdirs.
- In case there is no eval dataset, the experiment name must match a provider's name and this
  will be used to run the predictions.
- In case there is eval dataset, the experiment name is just a tag and the information about the
  prediction provider will be extracted by the corresponding column of the parquet.
- If there is not eval dataset and the experiment name does not match any prediction provider,
  an exception is raised.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: MultiEvalutor rename the GT_LEAF_DIR and introduce the EVALUATIONS_DIR to make the dir structure
created/used by MultiEvaluator the same with the ones created by the CLI

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Change the pipeline settings of Docling to use 16 CPU threads.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: MultiEvaluator improve logging

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Fix the MultiEvaluator.load_multi_evaluation() to properly scan the multi evalution dir structure

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Ensure to use all CPU cores for the DoclingPredictionProvider

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

---------

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>
2025-05-15 17:00:33 +02:00
Nikos LivathinosandChristoph Auer dee40e8f7d feat: Consolidate multiple evaluation results and generate a comparison matrix (#64)
* feat: Introduce the pred_modalities parameter in the BasePredictionProvider and its implementations

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Refactor main:get_prediction_provider() to add parameter that controls the visualizations.
Refactor the evaluate() to return the DatasetEvaluation as object.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce the MultiEvaluator that can generate ground truth and prediction datasets and also
compute the evalution across multiple providers and modalities. Add unit test.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Update toml dependencies to include pandas, openpyxl

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce staticmethod MultiEvaluator.load_multi_evaluation() to load multi-evaluations from
the disk. Update unit tests.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Allow PENDING in the _accepted_status of BaseEvaluator

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Modifications in the unit test of MultiEvaluator. Code clean up.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Introduce the Consolidator class that collects evaluation results and generates one excel report

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Improve the header names of excel export

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the DatasetLayoutEvaluation with DatasetStatistics for all metrics

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the Consolidator to include the standard deviation for each metric

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Update the pyproject.toml to pin to the  docling branch that supports the RT-DETRv2 model

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the DatasetEvaluation to contain the evaluated and rejected samples. The rejected ones
are itemized per rejection type.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Extend the Consolidator to include the samples (evaluated, rejected) in the generated excel

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Refactor the directory structure for MultiEvaluator and Consolidator classes.
- Refactor the generated excel matrix to include the experiment and provider columns.
- Refactor the BasePredictionProvider and all providers to have class attributes for the
  prediction_provider_type and prediction_modalities.
- Introduce CLI in the examples for the generation of the consolidation matrix.

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: Remove ConversionStatus.PENDING accepted status from BaseEvaluator

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* fix: MultiEvaluator fix the load_multi_evaluation()

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* feat: Add the class attributes for supported modalities in all prediction providers

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* chore: Fix code typos

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* Address predictor_info TODOs

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* fix: Add the test_multi_evaluator as a pytest dependency for test_consolidator

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>

* Repin to docling release

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* Regenerate lock file

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

* fix: Use defaultdict for rejected_samples

Signed-off-by: Christoph Auer <cau@zurich.ibm.com>

---------

Signed-off-by: Nikos Livathinos <nli@zurich.ibm.com>
Signed-off-by: Christoph Auer <cau@zurich.ibm.com>
Co-authored-by: Christoph Auer <cau@zurich.ibm.com>
2025-04-22 12:59:21 +02:00