Skip to content

AI provenance metadata in DOCX custom XML parts #1546

Description

@HMAKT99

Context

AI-generated Word documents are everywhere — reports, proposals, analyses created by LLMs. But the DOCX format carries no metadata about AI generation, trust level, or source provenance.

python-docx supports custom XML parts (customXml), which is the natural place to store this.

The Question

Is there a recommended pattern for storing custom provenance metadata in DOCX files via python-docx? Specifically:

  1. Can customXml parts be read/written through the current API?
  2. Is there interest in a helper for AI provenance (e.g., document.core_properties.ai_generated = True)?

Why This Matters

EU AI Act Article 50 (August 2, 2026) requires transparency metadata on AI-generated content. DOCX is a primary format for enterprise content. Having a standard way to mark AI provenance in Word docs is becoming a compliance requirement.

Existing Approach

AKF embeds provenance into DOCX custom XML:

<akf:metadata>{"v":"1.0","claims":[{"c":"Q3 report","t":0.85,"src":"SEC 10-Q"}]}</akf:metadata>

But curious if python-docx has its own approach or if this should be handled at the application level.

Activity

  1. HMAKT99 commented on Mar 28, 2026

    @HMAKT99
    Author

    Created a working example: python_docx_provenance.py — shows embedding AI provenance in DOCX custom XML using python-docx + AKF.

  2. scanny commented on May 18, 2026

    @scanny
    Contributor

    @HMAKT99 this is a very interesting question and I'm sure is of growing importance.

    Off the top of my head I'd be strongly inclined to use one of the properties mechanisms rather than CustomXml if it would get the job done.

    There is the notion of:

    • core-properties (docProps/core.xml): "core" as in "Dublin-Core", like standard library and information-science classifiers such as created on, author, etc. This is a fixed schema of properties.
    • app-properties (docProps/app.xml), perhaps known as extended properties sometimes: These are things like company, manager, editing time, page count etc. This is also a fixed schema and not yet implemented in python-docx.
    • custom-properties (docProps/custom.xml) (not to be confused with CustomXml functionality associated with structured-data tags (SDT). This schema is wide open so seems like the best fit off the top of my head.

    It seems like it would be important for folks to use a shared schema for AI generation metadata. It's not going to be much good if readers don't know where to find the metadata and how to interpret it.

    It seems like these folks have something going in that regard: https://c2pa.org/

    Looks like what they propose can't fit in a key-value pair though, so that would require a CustomXml part it looks like. Also, what they're up to seems to be pretty strongly oriented toward digital media like images and video, so I'm not entirely sure how that would apply to a generated document.

    Regular authorship attribution could be accomplished with core-properties for author and description. And it seems like custom-properties could add a lot too, just not things like a cryptographic signature perhaps, like you might use to determine whether something had been modified since the metadata was applied.

    What's your point of view on this?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions