Qianfan-OCR: A Unified End-to-End Model for Document Intelligence

AK's Picks Multimodal Intelligence Vision-Language Pre-training VRDU CV FEOE
我们推出了“千帆OCR”(Qianfan-OCR),这是一款参数量达40亿的端到端视觉—语言模型,首次在单一架构内统一实现了文档解析、版面分析与文档理解三大核心能力。该模型可直接将文档图像转换为结构清晰的Markdown格式,并支持多种基于提示词驱动的任务,包括表格提取、图表理解、文档问答以及关键信息抽取。针对端到端OCR模型普遍缺失显式版面分析能力的问题,我们提出了“版面即思维”(Layout-as-Thought)机制:当模型识别到特定的“思考令牌”(think tokens)时,将自动触发一个可选的推理阶段,先生成结构化的版面表征——包括各元素的边界框(bounding boxes)、元素类型(element types)及阅读顺序(reading order),再生成最终输出;此举不仅重建了模型对物理版面的感知与定位能力,还显著提升了其在复杂版式文档上的处理精度。在权威评测基准OmniDocBench v1.5与OlmOCR Bench上,“千帆OCR”分别以93.12分和79.8分的成绩位居所有端到端模型榜首;在OCRBench、CCOCR、DocVQA及ChartQA等主流评测中,其表现亦可媲美同规模通用视觉语言模型(VLMs);而在公开的关键信息抽取基准测试综合平均分上,更以领先优势超越Gemini-3.1-Pro、Seed-2.0及Qwen3-VL-235B等前沿大模型,位居第一。该模型已通过百度智能云千帆大模型平台向公众开放使用。
We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conversion and supports diverse prompt-driven tasks including table extraction, chart understanding, document QA, and key information extraction. To address the loss of explicit layout analysis in end-to-end OCR, we propose Layout-as-Thought, an optional thinking phase triggered by special think tokens that generates structured layout representations -- bounding boxes, element types, and reading order -- before producing final outputs, recovering layout grounding capabilities while improving accuracy on complex layouts. Qianfan-OCR ranks first among end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8), achieves competitive results on OCRBench, CCOCR, DocVQA, and ChartQA against general VLMs of comparable scale, and attains the highest average score on public key information extraction benchmarks, surpassing Gemini-3.1-Pro, Seed-2.0, and Qwen3-VL-235B. The model is publicly accessible via the Baidu AI Cloud Qianfan platform.
许愿