Devesh JoshiCo-founder, product
Nine years building AI platforms serving 12,000+ engineers. LLM platforms, agentic systems (MCP), and multi-model safety evaluation.
Multimodal AI in invoice processing: Extracting tables without brittle OCR templates
Replacing fragile bounding-box OCR templates with vision-language models for instant, zero-shot PDF invoice and contract data extraction.
The fragility of legacy OCR templates
Traditional optical character recognition (OCR) relies on rigid coordinate templates. When a vendor modifies line padding or table column order, parsing breaks.
Vision LLMs interpret documents visually, comprehending tables, footers, and line items regardless of template layout shifts.
Topic Focus & Target Concepts
This is the work behind our AI and automation practice — agents with real grounding, voice intake, and retrieval that answers from your records rather than the model's training data.
Agents with real tool accessRunning into this in your own stack? Twenty minutes, no deck.
Book the call