Automated Table Extraction from Scanned Oil & Gas Documents

What do you do when every off-the-shelf model fails on your scans, and the data still has to come out of the tables?

A large Middle Eastern oil and gas operator was processing PDFs, technical drawings, and scanned documents in a semi-manual way. Azati ran a full R&D cycle, from task setup and data labeling to models trained from scratch and production deployment, and replaced that process with a fully automated tabular-data extraction pipeline.

Automate my document data extraction
Fully automated

tabular-data extraction pipeline, replacing semi-manual document processing

12 months

from task setup and data labeling to production deployment

From scratch

table detection and OCR models, trained in-house after off-the-shelf models fell short

What problem did this platform solve?

It replaced semi-manual handling of tables locked inside scans and PDFs with an automated pipeline that outputs structured data.

An oil and gas operator runs on documentation: PDFs, technical drawings, and scanned paper records, the kind of unstructured engineering data that is hard to search and reuse. The facts that matter often sit in tables, and a scan gives you a picture of a table, not data. Before this project, getting those values out was a semi-manual job.

The client needed a platform that could take these document types and return structured tabular data. Azati was brought in on the strength of its computer vision and document intelligence work, and the engagement ran for 12 months as a Dedicated Team.

Technologies used

Python
Python
PyTorch
PyTorch
PyTorch Lightning
PyTorch Lightning
NumPy
NumPy
Albumentations
Albumentations
OpenCV
OpenCV

How does the pipeline turn a scanned page into structured data?

ModuleWhat it does
Table detectorFinds tables on scanned images.
OCR moduleExtracts the text from the detected table regions.
Table structure detectionRecovers the structure of each table by segmentation, so layouts that are not regular grids can be read.
Post-processingTurns recognized text and detected structure into structured tabular data.
Integration APIExposes the pipeline's results for integration with other systems.

What made this project harder than it looked?

The obvious route, plugging in a pretrained model, did not work on this client's documents, and every alternative needed its own experiments.

Challenge 01

Ready-made models were not good enough

The team started by reviewing existing solutions, and the answer was clear quickly:

  • Many off-the-shelf models showed weak extraction quality on the client's documents.
  • The team reviewed existing solutions before choosing its own approach.
  • The weak results pushed the team toward its own research and training from scratch.
#1
Challenge 02

Labeling was an open question

Training from scratch needs labeled data, and how you label decides what the model can learn:

  • The team tried many variants of document labeling before settling on what worked.
  • Labeling and training variants were compared across many experiments.
  • Data collection and labeling were part of the R&D cycle, alongside task setup and model training.
#2
Challenge 03

Three document types, one pipeline

The platform had to handle a range of inputs, not a single clean format:

  • PDFs, technical drawings, and scanned documents all had to feed the same extraction flow, a typical document digitization workflow problem.
  • The table detector was built to work on scanned images.
  • Drawing title blocks look like tables but are built from irregular blocks, which a column-based approach cannot handle.
#3

How did Azati get from failed models to production?

By treating it as research: test what exists, choose an approach, label data, train from scratch, and only then build the product around the models.

Reviewing existing solutions first

The team looked at what was already available before building anything. That review is what showed the ready-made options would not reach the quality the client needed.

Choosing table detection plus OCR

Azati split the task into two steps: find the table, then run OCR data extraction on it.

Treating table structure as segmentation

Existing solutions were largely built around column detectors, which would have limited the platform. Azati framed table structure detection as a segmentation task, an unusual choice that handles layouts which are not regular grids.

Labeling data and trying many variants

The team set up the task, collected the data, and labeled it, trying many labeling approaches along the way.

Training detection and OCR models from scratch

With a labeling scheme that worked, the team trained both models in-house. The toolchain included Python, PyTorch Lightning, and Albumentations.

Building post-processing and deploying to production

Raw OCR output is not yet data. The team wrote the post-processing that structures it, wrapped the pipeline in an integration API, and deployed it to production.

Are your tables still trapped in scans and PDFs?

Send us a few representative documents. We will tell you whether table detection plus OCR fits them, and what the first round of data labeling would need to look like.

Automate my document data extraction

What did Azati build?

Five modules that work as one pipeline: a table detector, an OCR module, table structure detection, a post-processing layer, and an integration API.

01

Table detector for scanned images

The detector locates tables on scanned pages. It was trained from scratch on data the team collected and labeled for this client.

Key capabilities:
  • Table detection on scanned images
  • Model trained from scratch on labeled client data
  • Built with PyTorch, Albumentations, and OpenCV
02

OCR module

The OCR module extracts text from detected tables. Like the detector, it was trained in-house rather than taken off the shelf.

Key capabilities:
  • Text extraction from detected table regions
  • OCR model trained from scratch
  • Trained with PyTorch and PyTorch Lightning
03

Table structure detection

Existing solutions were largely built around column detectors, which would have limited the platform. Azati approached table structure as a segmentation task instead, an unusual choice for this kind of solution. It made it possible to extract data even from scanned AutoCAD drawings, where the title block in the bottom right corner has a tabular structure but is really a tetris of blocks.

Key capabilities:
  • Trained with PyTorch, PyTorch Lightning and OpenCV from scratch
  • Table structure detection framed as a segmentation task
  • Not limited by the column-detector approach
  • Title block extraction from scanned AutoCAD drawings
04

Post-processing

This layer converts recognized text and detected structure into structured tabular data.

Key capabilities:
  • Structuring of raw OCR output into tables
  • NumPy in the toolchain
05

Integration API

An API for integration exposes the pipeline's results to other systems.

Key capabilities:
  • API for integration
  • Production deployment

Screenshots

Automated Table Extraction from Scanned Oil & Gas Documents

What changed for the client?

AreaBefore AzatiAfter Azati
Document processingSemi-manualFully automated pipeline
Tabular-data extractionDependent on a manual stepAutomated, from scan to structured data
Model qualityOff-the-shelf models with weak extraction qualityDetection and OCR models trained from scratch
ApproachExisting solutions reviewed firstTable detection plus OCR, chosen after that review

What did automation change in practice?

It removed the manual step from table extraction and gave the client a fully automated tabular-data pipeline.

The engagement did not report numeric benchmarks, so the honest summary is about process: what used to need a person now runs end to end.

The manual step is gone

Table data no longer depends on semi-manual processing. Documents enter the pipeline and structured data comes out.

One pipeline covers several document types

PDFs, technical drawings, and scanned documents are covered by one platform, including the title blocks of scanned AutoCAD drawings.

Models trained from scratch, not taken off the shelf

The detector and OCR were trained in-house on data the team collected and labeled, after ready-made models showed weak extraction quality.

An integration API is part of the platform

The platform exposes its results through an API for integration with other systems.

The relationship continued

The client kept working with Azati after this engagement.

What did the team learn from this project?

Three lessons: read the research before you buy the plugin, treat labeling as a modeling decision, and choose the approach after comparing, not before.

Reading papers beat waiting for a library

When ready-made models fell short, the team read scientific papers and ran its own research. That habit is the expertise that stays inside Azati after the project.

Labeling is a modeling decision, not a preparation step

The team tried many ways of labeling documents. Labeling was treated as part of the research, not as a task finished before it.

Choose the approach after comparing, not before

The team compared existing solutions and approaches before settling on table detection plus OCR. The approach was the outcome of that comparison.

What should you check before choosing a table extraction approach?

What to checkWhy it mattersWhen it becomes urgent
Pretrained models on your own scansBenchmark quality often does not hold on real documents, and this is cheap to test.Before any decision to build or buy.
Labeling approachThe way you label sets the ceiling for what a trained model can learn.When off-the-shelf models fall short and training becomes necessary.
Detection separate from OCRTwo focused models are easier to measure and improve than one that does everything.When documents mix tables with drawings or other content.
Output and integrationExtracted data only matters if it reaches the systems that use it.When the pipeline moves from a pilot to production.

Team composition

Azati ran the project as a Dedicated Team led by team leads.

  • Team leads who led the team through the full R&D cycle: task setup, data collection and labeling, comparison of approaches, training of the table detection and OCR models from scratch, post-processing, and production deployment.

How was the engagement delivered?

Dedicated Team over 12 months

A Dedicated Team suits research-heavy work, where the route to production is found through experiments.

Full R&D cycle in one team

The same Azati team handled task setup, data labeling, model training, post-processing, and deployment.

Delivered without third-party vendors

Azati delivered the project with its own engineers from start to finish.

Who this approach is relevant for

It fits organizations whose key data sits in tables inside scans, PDFs, or technical drawings, and where off-the-shelf OCR does not reach usable quality.

  • Oil and gas, energy, and engineering companies with large archives of scanned documentation
  • Teams whose current process for extracting table data is manual or semi-manual
  • Utilities and energy teams in the Middle East planning AI for engineering documentation
  • Engineering organizations that need document AI for regulated engineering workflows
  • Companies that tried generic OCR or pretrained models and saw weak extraction quality
  • Organizations ready to invest in custom model training on their own labeled data
  • Teams that need extracted data delivered through an API into existing systems
  • Programs that benefit from a document digitization strategy before any tooling is chosen

Frequently asked questions

If you're the one who has to explain to a CFO or operations head why a generic OCR tool is not enough for your scans, this FAQ is written for that conversation.

Automatic extraction works in stages: detect the tables, read their text with OCR, then structure the result.

Azati built exactly this pipeline for an oil and gas operator, with a table detector on scanned images, an OCR module, table structure detection, post-processing, and an API for integration.

Azati trained its own models because many ready-made models showed weak extraction quality on the client's documents.

After reviewing existing solutions, the team chose table detection plus OCR and trained both from scratch on labeled client data.

The platform extracts data from a range of document types, including PDFs, technical drawings, and scanned documents.

Its table detector works on scanned images, and the OCR module reads the detected tables.

A full R&D cycle covers task setup, data collection and labeling, comparison of approaches, model training, post-processing, and production deployment.

Azati's team carried this project through every one of these steps in 12 months.

The client moved from semi-manual document processing to a fully automated pipeline for extracting tabular data.

The engagement did not publish numeric benchmarks, so Azati describes the result as a change in process: the manual step is gone, and extraction runs end to end.

A title block looks like a table but is built from irregular blocks, so column detection does not fit it.

Azati's platform treats table structure detection as a segmentation task, which lets it extract data from the title block in the bottom right corner of scanned AutoCAD drawings.

Azati built the platform in Python with PyTorch, PyTorch Lightning, NumPy, Albumentations, and OpenCV.

The models for table detection and OCR were trained in-house, and the platform exposes its results through an integration API.

Azati delivered the project as a Dedicated Team over 12 months, without third-party vendors.

The client continued working with Azati after this engagement.

Have documents your current OCR cannot read into tables?

Bring a sample of your hardest scans. We will review them with you and outline what a detection plus OCR pipeline would take on your data.

Talk to our document AI team

Azati's related document AI and engineering data extraction expertise

Explore our successful projects and see how Azati delivers measurable results for our clients.

Automated Tag Extraction and AVEVA Table Generation from Engineering Drawings

Automated Tag Extraction and AVEVA Table Generation from Engineering Drawings

300 engineering documents processed per hour by the extraction pipeline
3 months to deliver the full MLOps pipeline from data annotation to production
Multi-level tag classification using visual context from drawings, not template matching
  • Python
  • PyTorch
  • OpenCV
  • FastAPI
  • AWS
  • Docker

⚡ Pain Points We Tackled

A large oil and gas engineering operator needed tag data extracted and classified from engineering drawings, categorized for AVEVA Engineering tables, and equipment connections identified from wiring diagrams. The manual process required multiple engineers and could not keep pace with document volume.

Our Approach

Azati built a full pipeline from data annotation to production that extracts and classifies tags, identifies equipment connections, and generates structured AVEVA Engineering tables, delivered in three months.

Applied Methods and Practices

  • Tag extraction and classification: Identifies tags on engineering drawings and classifies them using visual context.
  • Equipment connection identification: Detects connections between equipment from wiring diagrams.
  • AVEVA table generation: Produces structured tables ready for AVEVA Engineering.

Solution Features

  • 300 documents per hour: The pipeline processes up to 300 engineering documents per hour.
  • Three months to production: The full MLOps pipeline, from data annotation to deployment, was delivered in three months.
Multi-Document Engineering Drawing AI: P&ID, Electrical, HVAC & Fire/Gas

Multi-Document Engineering Drawing AI: P&ID, Electrical, HVAC & Fire/Gas

4 structurally different drawing types recognized by one pipeline: P&ID, fire and gas, HVAC, electrical
15 months from first annotation task to a working production pipeline
Dedicated team spanning development and data annotation
  • Python
  • PyTorch
  • PyTorch Lightning
  • Ultralytics YOLO
  • OpenCV
  • FastAPI

⚡ Pain Points We Tackled

A refining and petrochemical operator was extracting equipment tags and attributes from engineering drawings by hand, across four structurally different document types: P&ID, fire and gas loop diagrams, HVAC schematics, and electrical one-line diagrams.

Our Approach

Azati built a computer vision and OCR pipeline that detects objects, matches tags, classifies equipment, and extracts attributes automatically, packaging the result for direct upload into the client's engineering data platform.

Applied Methods and Practices

  • Object detection across drawing types: Recognizes instruments and equipment in four structurally different document types.
  • Tag matching and OCR: Reads tags and matches them to detected objects.
  • Attribute extraction: Pulls equipment attributes into structured output ready for upload.

Solution Features

  • One pipeline, four drawing types: P&ID, fire and gas, HVAC, and electrical diagrams are handled by a single pipeline.
  • 15 months to production: The Dedicated Team took the work from the first annotation task to a working production pipeline.
AI-Powered PEFS Digitization and DEXPI Conversion

AI-Powered PEFS Digitization and DEXPI Conversion

35,000 PEFS files in the archive, previously unsearchable
60-80% estimated reduction in manual work based on pilot results
DEXPI industry standard adopted for all converted output, enabling downstream system integration
  • Python
  • PyTorch
  • OpenCV
  • FastAPI
  • MongoDB
  • Oracle Cloud Infrastructure

⚡ Pain Points We Tackled

A large Middle Eastern oil and gas operator held 35,000 PEFS files across AutoCAD, PDF, TIFF, and JPEG formats, with no way to search them by equipment or process element.

Our Approach

Azati built an AI-powered pipeline that extracts and classifies data from these drawings, converts them to the DEXPI industry standard via Proteus XML, validates the output, and generates SVG visualizations.

Applied Methods and Practices

  • Drawing data extraction and classification: Extracts and classifies data from drawings across multiple file formats.
  • DEXPI conversion: Converts drawings to the DEXPI industry standard via Proteus XML and validates the output.
  • SVG visualization: Generates SVG visualizations of the converted drawings.

Solution Features

  • 35,000 files made queryable: A static archive became an intelligent, queryable engineering data layer.
  • 60-80% less manual work: Estimated from pilot results, with all converted output following the DEXPI standard.

Last updated

Got a job for Azati? Let’s talk business!

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

What's next?

  • 1. Tell Us Your Story
    Describe your project. We come back within 24 hours with team availability and a rough plan. NDA on request before the first call.
  • 2. Get Your Roadmap
    Receive a detailed proposal with scope, team composition, timeline, and costs tailored to your goals.
  • 3. Start Building
    Azati aligns on details, finalize terms, and launch your project with full transparency.