Skip to main content
Microsoft Word is a word processor developed by Microsoft.
本文介绍如何load Word documents into a document format that we can use downstream.

Using Docx2txt

Load .docx using Docx2txt into a document.

Using unstructured

Please see Unstructured for more instructions on setting up Unstructured locally, including setting up required system dependencies.

Retain elements

在底层,Unstructured 为不同的文本块创建不同的 “elements”。默认情况下我们将它们组合在一起,但你可以通过指定 mode="elements" 轻松保持这种分离。

Using Azure AI document intelligence

Azure AI Document Intelligence (formerly known as Azure Form Recognizer) is machine-learning based service that extracts texts (including handwriting), tables, document structures (e.g., titles, section headings, etc.) and key-value-pairs from digital or scanned PDFs, images, Office and HTML files. Document Intelligence supports PDF, JPEG/JPG, PNG, BMP, TIFF, HEIF, DOCX, XLSX, PPTX and HTML.
This current implementation of a loader using Document Intelligence can incorporate content page-wise and turn it into LangChain documents. The default output format is markdown, which can be easily chained with MarkdownHeaderTextSplitter for semantic document chunking. You can also use mode="single" or mode="page" to return pure texts in a single page or document split by page.

Prerequisite

An Azure AI Document Intelligence resource in one of the 3 preview regions: East US, West US2, West Europe - follow this document to create one if you don’t have. You will be passing <endpoint> and <key> as parameters to the loader. pip install -qU langchain langchain-community azure-ai-documentintelligence