Skip to content

TIKA-4831: Add content-based detection and a parser for GeoGebra files (ggb, ggs, ggt) - #3044

Open
dschmidt wants to merge 2 commits into
apache:mainfrom
dschmidt:feat/geogebra
Open

TIKA-4831: Add content-based detection and a parser for GeoGebra files (ggb, ggs, ggt)#3044
dschmidt wants to merge 2 commits into
apache:mainfrom
dschmidt:feat/geogebra

Conversation

@dschmidt

Copy link
Copy Markdown
Contributor

Adds detection and parsing support for the GeoGebra file formats.

Issue: https://issues.apache.org/jira/browse/TIKA-4831

Detection

  • application/vnd.geogebra.file (*.ggb) and application/vnd.geogebra.tool (*.ggt) are now sub-class-of application/zip, so the filename hint survives magic detection instead of being discarded in favor of plain application/zip
  • new mime types: application/vnd.geogebra.slides (*.ggs, zip-based) and application/vnd.geogebra.pinboard (*.ggp, JSON-based)
  • a new GeoGebraDetector (zip container detector, ZipFile and streaming mode) recognizes the formats without a filename by their well-known entries: geogebra.xml (worksheet), structure.json + _slideN/geogebra.xml (Notes/Slides), geogebra_macro.xml (tool). Since a worksheet with macros contains both geogebra.xml and geogebra_macro.xml, the decision is made after all entry names have been seen
  • ZipParser.ZIP_SPECIALIZATIONS is kept in sync with the new registry entries

Parser

A new GeoGebraParser (miscoffice module) for ggb/ggs/ggt:

  • metadata: construction title/author/date become dc:title/dc:creator/geogebra:date, plus geogebra:appName, geogebra:appVersion, geogebra:formatVersion, geogebra:id; Notes/Slides set xmpTPg:NPages
  • content: user-visible text as XHTML paragraphs - string-literal expression elements of text objects, rich-text content runs (JSON), element captions, macro names and help texts; Notes/Slides emit one div per slide in structure.json order
  • the representative rendering (geogebra_thumbnail.png at the root, or the first slide's thumbnail) is emitted as an embedded document marked embeddedResourceType=THUMBNAIL, following the existing convention in the OOXML, ODF and iWork parsers, so /unpack/all sidecars identify the preview image; other slides' thumbnails are redundant renderings and are skipped
  • any other embedded file (e.g. inserted pictures) is emitted as an embedded document

XML is parsed through XMLReaderUtils.parseSAX like the other parsers in the module; structure.json and rich-text runs use Jackson (new jackson-databind dependency in the miscoffice module, version managed by the existing BOM import).

Testing

  • new unit tests: GeoGebraDetectionTest (zip-commons) and GeoGebraParserTest (miscoffice) with crafted ggb/ggs/ggt fixtures; the ggs fixture deliberately orders _slide1 before _slide0 in structure.json to pin the ordering behavior
  • full test suites of tika-core, tika-parser-zip-commons, tika-parser-miscoffice-module and tika-parser-pkg-module pass, plus CompositeZipContainerDetectorTest in the integration tests
  • verified end-to-end against a real-world GeoGebra Notes file: detected as vnd.geogebra.slides, text extracted, thumbnail emitted with the THUMBNAIL marker

Mark application/vnd.geogebra.file and .tool as sub-classes of
application/zip so the filename hint survives magic detection, and add
the (IANA-registered) application/vnd.geogebra.slides (*.ggs) and
application/vnd.geogebra.pinboard (*.ggp) types.

Add a GeoGebraDetector to the zip container detectors that recognizes
the formats without a filename by their well-known entries: geogebra.xml
(worksheet), structure.json plus _slideN/geogebra.xml (Notes/Slides) and
geogebra_macro.xml (tool), in ZipFile and streaming mode.

Keep ZipParser.ZIP_SPECIALIZATIONS in sync with the new registry
entries.
Parse geogebra.xml/geogebra_macro.xml with the pooled, hardened SAX
path: construction title/author/date and the app name/version/format/id
become metadata, and the user-visible text (string-literal expressions,
rich-text content runs, captions, macro names and help texts) is emitted
as XHTML paragraphs. Notes/Slides files emit one div per slide in
structure.json order and set xmpTPg:NPages.

The representative rendering - geogebra_thumbnail.png at the root, or
the first slide's thumbnail - is emitted as an embedded document marked
embeddedResourceType=THUMBNAIL so unpack sidecars identify the preview
image; other slides' thumbnails are redundant renderings and are
skipped. Any other embedded file (e.g. inserted pictures) is emitted as
an embedded document.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant