Skip to main content
Microsoft Word is a word processor developed by Microsoft.
This covers how to load Word documents into a document format that we can use downstream.

Using Docx2txt

Load .docx using Docx2txt into a document.

Using Unstructured

Please see this guide for more instructions on setting up Unstructured locally, including setting up required system dependencies.

Retain Elements

Under the hood, Unstructured creates different โ€œelementsโ€ for different chunks of text. By default we combine those together, but you can easily keep that separation by specifying mode="elements".

Using Azure AI Document Intelligence

Azure AI Document Intelligence (formerly known as Azure Form Recognizer) is machine-learning based service that extracts texts (including handwriting), tables, document structures (e.g., titles, section headings, etc.) and key-value-pairs from digital or scanned PDFs, images, Office and HTML files. Document Intelligence supports PDF, JPEG/JPG, PNG, BMP, TIFF, HEIF, DOCX, XLSX, PPTX and HTML.
This current implementation of a loader using Document Intelligence can incorporate content page-wise and turn it into LangChain documents. The default output format is markdown, which can be easily chained with MarkdownHeaderTextSplitter for semantic document chunking. You can also use mode="single" or mode="page" to return pure texts in a single page or document split by page.

Prerequisite

An Azure AI Document Intelligence resource in one of the 3 preview regions: East US, West US2, West Europe - follow this document to create one if you donโ€™t have. You will be passing <endpoint> and <key> as parameters to the loader. pip install -qU langchain langchain-community azure-ai-documentintelligence

Connect these docs programmatically to Claude, VSCode, and more via MCP for real-time answers.