* feat(vtt): export and save to WebVTT format
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* fix(vtt): omit empty blocks in parsing
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* test(vtt): add tests for exporting and saving to WebVTT
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
PDF hyperlinks may contain relative paths, internal bookmarks, or
fragment-only references that are not valid absolute URLs. The strict
AnyUrl validation on PdfHyperlink.uri caused the entire page preprocess
stage to fail when such URIs were encountered, resulting in empty
documents and lost content.
Change uri type to Union[AnyUrl, str] with a field_validator that
attempts AnyUrl parsing first (preserving structured metadata like
scheme/host/path) and falls back to str for non-absolute URIs.
Signed-off-by: Ultizan <ultizan@gmail.com>
* fix(chunker): propagate 'traverse_pictures' parameter from serializer to chunker
Propagate 'traverse_pictures' parameter from the 'serializer_provider' to the HierarchicalChunker, to ensure
that chunks include any TextItem children from PictureItem.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* test(chunker): simplify imports
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* fix(serializer): add 'traverse_pictures' to CommonParams
Expose 'traverse_pictures' parameter to serialier common parameters to control
whether to traverse into PictureItem objects to serialize their children.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* test(serializer): add test for 'traverse_pictures' parameter
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* fix(markdown): add an option to compact table serialization
Add an option in MarkdownTableSerializer to remove padding in tables in markdown serialization.
Propagate the option to MarkdownDocSerializer and DoclingDocument.
Remove unnecessary use of 'mode' argument in 'open' function when mode is 'r'.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(markdown): keep alignment marks in '_compact_table'
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor: move WebVTT data model from docling
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* fix(webvtt): deal with HTML entities in cue text spans
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): support more WebVTT models
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(DoclingDocument): create a new provenance model for media file types
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): make WebVTTTimestamp public
Since WebVTTTimestamp is used in DoclingDocument, the class should be public.
Strengthen validation of cue language start tag annotation.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): set languages to a list of strings in ProvenanceTrack
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* tests(webvtt): add test for ProvenanceTrack
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): make all WebVTT classes public for reuse
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(webvtt): preserve newlines as WebVTTLineTerminator
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): set ProvenanceTrack time fields as float
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(webvtt): ensure start time offsets are in sequence
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(webvtt): improve regex to remove note,region,style blocks
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(webvtt): parse the WebVTT file title
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(webvtt): rebase to latest changes in idoctags
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* feat(webvtt): add WebVTT serializer
Add a DoclingDocument serializer to WebVTT format.
Improve WebVTT data model.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* fix(webvtt): add 'text/vtt' as extra mimetype
Add 'text/vtt' as extra MIME type to support WebVTT serialization, since it is not
supported by 'mimetypes' with python < 3.11
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): roll back DocItem.prov as list of ProvenanceItem
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* tests(webvtt): fix test with STYLE and NOTE blocks
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style(webvtt): apply X | Y annotation instead of Optional, Union
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): simplify TrackProvenance model with tags
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* refactor(webvtt): align class and field names to new 'source' type
Classes and fields that are related to the new source type should aign with their names.
The term 'provenance' will identify the legacy implementation.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* chore(DoclingDocument): drop the validation on field assignment
Drop the validation on field assignment in NodeItem objects.
Add the 'source' argument in the convenient function 'add_text' to create TextItem with track source data.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
refactor(webvtt): drop cue span classes, 'lang' and 'c' tags
Drop WebVTT formatting features not covered by Docling across formats.
Only 'u', 'b', 'i', and 'v' are supported and without classes.
Make 'v' tag explicit as 'voice' feature in SourceTrack class.
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* Added ruff to dev dependencies
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Added ruff settings to pyproject.toml as in docling
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Cleanup uf pyproject.toml
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Copied settings for ruff pre-commit hooks from docling
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Excluded test/data/** from ruff formatting / linting
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* ruff format
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Added some ignore statements to pyproject.toml such that ruff check raises fewer issues
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* ruff check --fix
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Ignored some more rules
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Fixed the rest of the errors that would only concern 1 - 3 files
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Added another ignore related to df for DataFrame names
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Modified CONTRIBUTING.md such that black / isort are replaced by ruff
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Added UP045 to ignore list such that Optional[...] does not raise
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Moved .flake8 configs to pyproject.toml
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Moved autoflake to be used with ruff
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Moved all .flake8 settings to pyproject.toml to be compatible with ruff (i.e. no separate [tool.flake8] section
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Removed flake8 from .pre-commit hooks
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Applied ruff format (again); formatted some files as the line-length = 120 equals now what was set for the .flake8 settings
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Set max-complexity to 30 (as was originally) in the pyproject.toml as one linting check would fail
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Adding PD901 to ignore list such that pre-commit hooks run fully again
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* Replaced dtype | None syntax by Optional[dtype] in remaining places
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
* chore: fix 'test' ref in pyproject
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove typing List, Set, Tuple, Dict
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove UP015 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove UP034 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: normalize dashes in comments and docstrings
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove PD901 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove C403 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove C403, C413, C416 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* style: remove E203, F811 check from ignore list
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
---------
Signed-off-by: Florian Schwarb <florian.schwarb@gmail.com>
Signed-off-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
Co-authored-by: Florian Schwarb <florian.schwarb@gmail.com>
Co-authored-by: Cesar Berrospi Ramis <ceb@zurich.ibm.com>
* feat(DocItem): Add comments field for linking annotations to document items
Implements support for linking comments (from Word/PPT documents) to their
annotated content using the established FloatingItem/RefItem pattern.
Changes:
- Add `comments: List[RefItem]` field to DocItem class
- Update `_update_breadth_first_with_lookup()` to handle comment references on deletion
- Bump CURRENT_VERSION to 1.9.0
- Fix version comparison bug (string vs integer for minor version)
- Add 4 new tests for comments functionality
- Update test data files for new schema
Closes: docling-project/docling#464
Related: docling-project/docling#2834
Signed-off-by: s1v4-d <leelasaisivasubrahmanyamdurga@gmail.com>
* improve comment Pydantic serialization
Signed-off-by: Panos Vagenas <pva@zurich.ibm.com>
* add add_comment, update tests
Signed-off-by: Panos Vagenas <pva@zurich.ibm.com>
* introduce fine-granular references with span ranges
Signed-off-by: Panos Vagenas <pva@zurich.ibm.com>
* simplify last test
Signed-off-by: Panos Vagenas <pva@zurich.ibm.com>
---------
Signed-off-by: s1v4-d <leelasaisivasubrahmanyamdurga@gmail.com>
Signed-off-by: Panos Vagenas <pva@zurich.ibm.com>
Co-authored-by: Panos Vagenas <pva@zurich.ibm.com>
* Working on ISO standard for IDocTags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the pre-commit, still tons of work to do
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* updated idoctag token class
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* need to start fixing the bugs and add proper tests for idoctags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the tests
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* refactored the testing for dcotags and idoctags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* updating the location tokens in idoctags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* removed the DocumentToken
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the idt quot issue
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* reformatted
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* Added IDocTagsSerializationMode
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* working on the deserializer
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the captions for floating items
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* still need to fix some seserialization tests
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* finally all is working and reformatted
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* made the OTSL serialization and deserialization self-contained in IDocTags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed some unclean code
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* did some fixes and clean up
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added extra tests, expanded to deserialize to deal with nested lists
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added the deserialization for formatting
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the formatting and the complex inline groups in nested lists
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added get_category
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* make IS_SELFCLOSING a set
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* removed the regex
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the if ... return into if ... elif ... return
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the facets parsing
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the footnotes for Tables and Pictures in IDocTags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* updated with latest doctags token table
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* removed private name
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
---------
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* feat: updating the idoctags serializer
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* feat: adding IDocTagsToken class
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fix the get_special_tokens
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fix the get_special_tokens (2)
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the doctags serializer: if the no_content flag is true, no double list_item, no content in tables and captions
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the list serializer in the doctags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added missing document serializer
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* removed the prints
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added tests for content suppression
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added extra tests for serialization
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* added extra tests
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* ran pre-commit
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* deal with differences between xml serialization in mac and linux
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* revert to old doctags serialization
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* revert to old doctags serialization (2)
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* revert to old doctags serialization (3)
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* revert to old doctags serialization (4)
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixing the tests
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed all the tests, except the no-content with lists
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* fixed the no-content in lists with doctags
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
* removed the prints
Signed-off-by: Peter Staar <taa@zurich.ibm.com>
---------
Signed-off-by: Peter Staar <taa@zurich.ibm.com>