overview
What is Cohere Parse 5?
Cohere Parse 5 is a document vision parsing AI tool developed by Cohere that enables enterprises to transform unstructured data contained within images and documents into structured, machine-readable data. It processes document pages as base64-encoded images (PDF, PowerPoint, or JPEG) through a single vision-language model pass to return structured Markdown, collapsing the traditional OCR-plus-model pipeline for efficiency. The model extracts text in reading order, tables (rendered as HTML), lists, forms, key-value pairs, images, captions, and bounding box coordinates for visual elements, supporting traceability for AI agents. This 2.3-billion-parameter vision language model is built on Cohere Labs' North-Micro-Vision-Instruct architecture, featuring an 8,192-token context window and a roughly 4.6-gigabyte footprint.
