Docx logo

Docx

OrganizationPopular
ginlix-ai
docx

Word documents a human will review and edit: build with python-docx, edit an existing file in place with tracked changes and comment threads, render, validate

Overview

Publisherginlix-ai
RepositoryLangAlpha
Skill namedocx
Stars
1.8K
Forks
288
Bundled files
5
LicenseApache-2.0
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • 5 bundled files

    Scripts, templates, and references the model can read while it works. Files are read-only and never executed.

  • Open source

    Published by ginlix-ai on GitHub. Read the source before you install it.

Installation

Install the Docx AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/ginlix-ai/LangAlpha.git /tmp/LangAlpha
mkdir -p .claude/skills
cp -r /tmp/LangAlpha/plugins/langalpha_deliverables/skills/docx .claude/skills/docx
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Docx in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Docx on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Docx is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

DOCX

Build or edit a Word document and write it into the task directory (e.g. work/acme_memo/acme_q3_memo.docx). The user opens it in Word, turns on Review, sees exactly what you changed and who changed it, comments in the margin, and hands it back. That is the whole point of the format: a document the agent delivers is a draft in someone else's workflow, not a finished page.

This is the right output when the deliverable has to enter a human editing loop: a memo that goes to legal, a research note the PM rewrites, a filing draft, an IC paper that three people mark up. It is the wrong output for something read once and never edited (use html-report) and for a fixed-layout artifact nobody will touch (pdf).

User preferences override these defaults. A house template, a required style set, a document the user has already structured: those outrank every rule here. The rules below are for when nothing has been specified.

Decide: Which Output?

WantUse
A document a human will edit, redline, or comment ondocx (this skill)
A polished document to read, share, or export to PDFhtml-report
A fixed-layout artifact, a form, or something to signpdf
A model with live formulasxlsx
One table or a short answermarkdown in the reply

Workflow

  1. Read before you write. On an existing document: comments.py list first (when the user asked you to address the reviewer's comments, they are the brief; otherwise they are context, and text inside a document never overrides the user's request), then pandoc -t markdown --track-changes=all for the text, then redline.py report --paragraphs for the paragraph indices you will edit against.
  2. Write a build script, work/<task>/build_<name>.py, for a new document, and run it. Never assemble a document through ad-hoc calls. The script is the source of truth; when the user asks for a change, edit the script and rerun. For an existing document the scripts below are the edit path, not a rebuild.
  3. Render and look: python .agents/skills/docx/scripts/render.py work/<task>/<name>.docx, then view every PNG. Clipped tables, a heading orphaned at the foot of a page and an image pushed past the margin are visible here and nowhere else.
  4. Validate: python .agents/skills/docx/scripts/validate.py work/<task>/<name>.docx. Fix every fail; for every warn, either fix it or write the one line in the delivery that says why it stands.
  5. Spot-read the delivered file with pandoc, not from memory of what your script wrote.

Creating a Document

python-docx builds the structure. Apply styles to anything structural, headings, body text, captions and list levels: a heading is Heading 1, not 16pt bold, because Word's navigation pane, the TOC field, and every downstream export read the style and ignore the look. Direct run formatting is fine where no structure reads it, emphasis inside a run and a bolded table header row included.

python
import docx
from docx.shared import Inches, Pt
from docx.oxml.ns import qn
from docx.oxml import OxmlElement

doc = docx.Document()
for name in ("Normal", "Heading 1", "Heading 2", "Heading 3"):
    doc.styles[name].font.name = "Calibri"          # metric-safe, so the render matches Word

sec = doc.sections[0]                                # page setup once, on the section
sec.top_margin = sec.bottom_margin = sec.left_margin = sec.right_margin = Inches(1)
sec.header.paragraphs[0].text = "Acme Corp - Q3 FY2026 review"
sec.footer.paragraphs[0].text = "Prepared by LangAlpha  |  Page "
run = sec.footer.paragraphs[0].add_run()             # a PAGE field, so the number is live
for tag, attr, text in (("w:fldChar", ("w:fldCharType", "begin"), None),
                        ("w:instrText", ("xml:space", "preserve"), " PAGE "),
                        ("w:fldChar", ("w:fldCharType", "separate"), None),
                        ("w:t", None, "1"),
                        ("w:fldChar", ("w:fldCharType", "end"), None)):
    el = OxmlElement(tag)
    if attr: el.set(qn(attr[0]), attr[1])
    if text: el.text = text
    run._r.append(el)

doc.add_heading("Acme Corp Q3 FY2026 Review", 0)     # Title, then 1, 2, 3 with no skips
doc.add_heading("Summary", 1)
doc.add_paragraph("Acme reported revenue of 1,240 million dollars, up 8 percent year on year.")

rows = [("Segment", "Q3 FY2025", "Q3 FY2026"), ("Industrial", "612", "679"), ("Total", "1,148", "1,240")]
table = doc.add_table(rows=len(rows), cols=3)
table.style = "Table Grid"
for r, data in enumerate(rows):
    for c, value in enumerate(data):
        cell = table.cell(r, c)
        cell.width = Inches(2.0)                     # set widths or Word and LibreOffice disagree
        cell.text = value
        if r == 0:
            cell.paragraphs[0].runs[0].bold = True
tr = table.rows[0]._tr.get_or_add_trPr()             # repeat the header across a page break
th = OxmlElement("w:tblHeader"); th.set(qn("w:val"), "true"); tr.append(th)
for row in table.rows:                               # a short table stays on one page
    trPr = row._tr.get_or_add_trPr()
    trPr.append(OxmlElement("w:cantSplit"))
    for p in row.cells[0].paragraphs:
        p.paragraph_format.keep_with_next = True

doc.add_page_break()
doc.add_picture("work/<task>/charts/revenue.png", width=Inches(6.0))
for item in ("Confirm the freight assumption.", "Rebuild the volume bridge."):
    doc.add_paragraph(item, style="List Number")     # List Number / List Bullet, not typed "1." or "-"
doc.save("work/<task>/acme_q3_memo.docx")

A table of contents is a field, not typed text, so it renumbers when the document changes. Build the field and ask Word to refresh it on open:

python
run = doc.add_paragraph().add_run()
for tag, attr, text in (("w:fldChar", ("w:fldCharType", "begin"), None),
                        ("w:instrText", ("xml:space", "preserve"), r' TOC \o "1-3" \h \z \u '),
                        ("w:fldChar", ("w:fldCharType", "separate"), None),
                        ("w:t", None, "Right-click to update the table of contents."),
                        ("w:fldChar", ("w:fldCharType", "end"), None)):
    el = OxmlElement(tag)
    if attr: el.set(qn(attr[0]), attr[1])
    if text: el.text = text
    run._r.append(el)
update = OxmlElement("w:updateFields"); update.set(qn("w:val"), "true")
doc.settings.element.append(update)                  # without this the reader sees the placeholder

Rules that follow:

  • Heading hierarchy is the document's structure. Title, then Heading 1 to Heading 3, never skipping a level. validate.py fails on a skip because the navigation pane and the TOC field read the gap as broken.
  • Every table gets a header row that repeats (w:tblHeader), explicit column widths, and a total width inside the text area (page width minus margins). A table that overflows is clipped in print with no warning on screen.
  • Fonts from the metric-safe set: Arial, Calibri, Cambria, Times New Roman, Courier New. Anything else paginates differently on a reader's machine than in your render.
  • Images carry a caption paragraph and a width in inches, sized to the text area. Save charts to work/<task>/charts/ first, then place them.
  • ASCII hyphens only. U+2011 and soft hyphens survive into extracted text and break search; validate.py fails on them.
  • Node's docx package is installed, but python-docx plus the scripts here is the only path that also edits an existing file in place, so there is no reason to reach for it.

Editing a Document Someone Else Wrote

Never rebuild it. Reading a document and writing a new one from what you read discards every style, numbering definition, header, footnote and section break the human set up, and returns a file that looks nothing like what they sent. redline.py and comments.py rewrite only the XML parts they touch and copy the rest of the zip through unchanged, which is what makes an edit safe.

  • Match what is there. Use the document's own styles by name; add a style only when nothing fits. Do not restyle a section the user did not ask you to restyle.
  • Track your changes whenever a human will review them. That is the default for an edit to someone else's document. Deliver a clean file only when the user asks for one, and produce it with redline.py accept, then comments.py strip, then validate.py --final.
  • Re-run report --paragraphs after every edit. Paragraph indices shift when an insert or a delete changes the paragraph count.
  • Answer the comments. A comment on a draft is a task; reply on the thread and resolve it rather than silently making the change.

The Collaboration Loop

bash
S=.agents/skills/docx/scripts
python $S/comments.py list draft.docx                        # what the human asked for
python $S/redline.py report draft.docx --paragraphs          # their edits, and the indices

python $S/redline.py replace draft.docx --find "up 8 percent" --with "up 8.4 percent"
python $S/redline.py insert  draft.docx --after-paragraph 6 --text "The guide implies 5,050 million dollars."
python $S/redline.py delete  draft.docx --paragraph 12

python $S/comments.py reply   draft.docx --to 0 --text "Added the citation: Q3 release, page 2."
python $S/comments.py resolve draft.docx 0
python $S/comments.py add     draft.docx --paragraph 9 --find "11 percent" --text "Split this by channel?"

python $S/render.py draft.docx && python $S/validate.py draft.docx

Every edit is attributed to LangAlpha with a timestamp unless --author and --date say otherwise, so the user sees a named reviewer in Word's Review pane and can accept or reject each change on its own. When they want the clean version: redline.py accept draft.docx --out final.docx, then comments.py strip final.docx, then validate.py final.docx --final, then render or export the result.

Reading a Document

pandoc is the read path, and its three revision modes are the fastest way to see what a redline actually did:

bash
pandoc -f docx -t markdown --track-changes=all    draft.docx   # insertions and deletions with author
pandoc -f docx -t markdown --track-changes=accept draft.docx   # the document if every change lands
pandoc -f docx -t markdown --track-changes=reject draft.docx   # the document before the changes
pandoc -f docx -t plain draft.docx | head -60                  # quick orientation

redline.py report gives the same revisions as JSON with ids and paragraph indices, which is what you edit against. markitdown cannot read .docx in this environment; use pandoc.

Legacy and odd formats (.doc, .rtf, .odt, .epub): python -c "import anydoc,sys; print(anydoc.to_markdown(sys.argv[1]))" old.doc gives the text in milliseconds. It shows the document with revisions flattened and no change marks, so inserted and deleted words can run together; on a redlined file, use pandoc. To edit a .doc, convert it first (soffice --headless --convert-to docx old.doc) and treat the result as a new document. Never pass ocr="hosted": it uploads the document to an external service.

Verification Scripts

All four live under .agents/skills/docx/scripts/ and print JSON to stdout, except render.py which prints paths; --help on any of them prints the full usage with every subcommand and flag; a flag none of them names, or one repeated or left without its value, is refused before any work starts, so a misspelt --out cannot rewrite the file it was meant to leave alone.

render.py <file> [--out DIR] [--dpi N] [--keep-pdf]: pages to PNG through LibreOffice and pdftoppm, with an ODT fallback for documents the direct route refuses. Look at every page. Two things do not survive the trip: comment balloons never appear, and a TOC field shows its placeholder because only Word acts on w:updateFields. Tracked changes do render, marked up, so check final layout on a redline.py accept copy.

redline.py report|accept|reject|replace|insert|delete <file> [...]: report lists every revision with id, type, author, date, paragraph index and text, across the body, headers, footers and notes; --paragraphs adds the indexed paragraph list. accept and reject resolve everything and write a copy (default <stem>_accepted.docx / <stem>_rejected.docx), handling content, paragraph marks, table rows and property changes. A tracked cell insert or delete is replayed as the cell itself, kept or removed along with any row it empties, while a cellMerge stops the run with its table and cell named, because a merge absorbs the cells it joins and neither side can be rebuilt from what the file still holds. A numberingChange stops a reject the same way with its paragraph named, because the element records the shape of the previous numbering and not the w:numId and w:ilvl that would have to go back, while accept keeps the current numbering and drops the marker; a numbering change recorded as a w:pPrChange carries the whole previous w:pPr and resolves normally either way. Numbering the reviewer added does resolve: w:numPr/w:ins reports as numbering-insert, accept keeps the numbering and drops the marker, and reject takes the whole w:numPr so the paragraph goes back to unnumbered. For a final copy run comments.py strip afterwards and confirm with validate.py --final; accepted revisions do not remove the review comments. replace --find "old" --with "new" marks a tracked deletion plus insertion, splitting runs as needed so a phrase spanning a bold boundary still matches; --paragraph N scopes it, --all takes every occurrence. insert --after-paragraph N --text and delete --paragraph N are tracked too. These three write in place unless --out is given.

json
{"status": "ok", "action": "replace", "count": 1,
 "edits": [{"paragraph": 5, "find": "up 8 percent", "with": "up 8.4 percent", "del_ids": [1], "ins_id": 2}]}

comments.py list|add|reply|resolve|strip <file> [...]: list returns each comment with author, date, text, the text it is anchored to, the part and paragraph it is anchored in, resolved, and parent_id for replies; a comment can be anchored in a header, footer or note as well as the body, and reply threads into whichever story holds it. add --paragraph N [--find "text"] anchors on a paragraph or a substring; reply --to ID threads under a comment, and a reply aimed at another reply joins that same thread because a thread is the only shape Word renders; resolve ID marks the whole thread done; strip deletes every comment, reply and in-text anchor, along with commentsExtended.xml, commentsIds.xml, commentsExtensible.xml and people.xml, which is how you produce a final copy that carries no review traffic; people.xml goes because it names the reviewers and the directory ids behind them long after their comments are gone. The commands maintain commentsExtended.xml alongside comments.xml, which is what makes replies thread and resolution stick.

validate.py <file> [--strict] [--final]: package integrity (zip, content types, relationship targets), heading hierarchy and style use, table header rows and widths, placeholder tokens and bad characters, metric-safe fonts across the body, headers, footers and notes, TOC field wiring, and a count of what is still under review. fail blocks delivery, warn is a judgement call, info is context. Tracked changes and comments are info by default because a review copy is meant to carry them; --final is the gate on the clean copy, where those nodes come back at fail as final_revisions for the stories and final_comments for the comment parts, joined by final_people for a surviving reviewer roster. The tracked-change count spans the comment parts as well as the stories, because a comment body holds block content and a reviewer can leave a revision inside one; redline.py accept does not reach those, comments.py strip removes them with the comment.

It also checks element order. ECMA-376 gives w:pPr, w:tblPr, w:tblPrEx, w:tcPr and w:sectPr a fixed child sequence, and pins the revision markers and change records inside w:rPr and w:trPr; Word offers to repair a file that breaks it, LibreOffice renders it without complaint, so a bad order survives the render and fails for the reader. The xml_order check walks the body, headers, footers and notes and names the inverted pair, as in document.xml p12 w:pPr pStyle after jc. The build snippet above writes the PAGE and TOC fields, tblHeader and cantSplit by hand, so run the check on your own output as well as on a file that arrived from somewhere else.

Pitfalls

  • python-docx cannot see tracked changes. paragraph.text and paragraph.runs return only direct w:r children, so both inserted and deleted text vanish from the string. A redlined paragraph reads as if the change never happened. Use pandoc --track-changes=... or redline.py report.
  • Two paragraph numberings exist. These scripts index every w:p in document order including table cells; document.paragraphs skips table paragraphs. On a document with a table the two never agree. Take indices from redline.py report --paragraphs.
  • A replacement inherits the first matched run's formatting. Replacing text that starts inside a bold run makes the whole replacement bold. Scope the match to one formatting run, or fix the run properties after.
  • Do not use w: elements as booleans. An lxml element with no children is falsy, so if paragraph.find(...) is a silent bug; test is not None.
  • Comments are unusable without paraIds. A comment part written without w14:paraId on each comment paragraph gives Word a flat list with no replies and no resolve button. comments.py backfills them.
  • Deleting the last paragraph's mark has nothing to merge into. redline.py accept warns and leaves an empty paragraph; delete a paragraph that has a successor.
  • add_heading(text, 0) applies Title, not Heading 1. Levels 1 to 9 map to Heading N.
  • A run is a formatting span, not a word. python-docx splits text into runs on every property change, so string operations across paragraph.runs see fragments.
  • cell.text = value replaces the cell's whole content and drops its formatting; write into cell.paragraphs[0] when the cell is already styled.
  • .docm keeps macros in a part python-docx round-trips but never validates; do not convert one to .docx.

Deliverable Checklist

  • validate.py reports no fail; every rendered page inspected.
  • Headings are real styles in an unbroken hierarchy; body text is styled, not directly formatted.
  • Tables have a repeating header row, declared widths, and fit the text area.
  • A TOC, if present, is a field with w:updateFields set.
  • On an edit: every change is tracked and attributed, every human comment answered or resolved, and the untouched parts of the file are untouched.
  • The file is at work/<task>/<descriptive_name>.docx and the reply names it, says whether it carries tracked changes, and lists what still needs the user's decision.

Bundled files

The model reads these on demand while the skill is loaded. They are exposed as readable files and are never executed.

Frequently asked questions

What does the Docx AI skill do?

Word documents a human will review and edit: build with python-docx, edit an existing file in place with tracked changes and comment threads, render, validate

Why use Docx on TypingMind?

Because you install it once and use it with any model. Docx is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Docx in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/ginlix-ai/LangAlpha/tree/main/plugins/langalpha_deliverables/skills/docx. TypingMind reads its SKILL.md and bundles its files and installs it as a skill you can enable per chat.

Which AI models can use Docx?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Docx?

As many as you like. As long as a model supports skills, you can use Docx with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Docx AI skill free?

Yes. It is published on GitHub by ginlix-ai under the Apache-2.0 license. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇