Extract repeatable, schema-validated data from PDFs and other unstructured sources, with agent- and human-in-the-loop workflows, approvals and audit-grade governance. SED is a standalone product, and the same capability powers the SPICE platform.

Sign-off workflows, role-based reviews and full audit trails ensure only approved data is released.
Agent-driven extraction, backed by OCR, validation rules and feedback loops, drives continuous improvement.
Configurable agents can perform the review step under confidence thresholds, with human override and final sign-off.
Receive files via SFTP, cloud storage, UI or API, validated and normalised automatically.
Agents leverage OCR to extract content and map it to your schema.
Validate, approve and publish with a complete audit trail.

SFTP, cloud object storage (S3/Blob/GCS), UI upload and REST endpoints simplify integration.
Agents leverage OCR to extract content and map it to your schema, including scanned or image-only PDFs.
Robust parsing targets complex tables with cell-level fidelity, minimising manual cleanup.
Map extracted fields to consistent, versioned schemas for analytics and APIs.
Configurable checks catch anomalies and route exceptions for review.
SSO (SAML/OIDC), RBAC, MFA, encryption in transit and at rest, and detailed audit logs.
Extract structured data from announcements, filings and reports.
Standardise submissions and evidence at scale.
Turn disclosure documents into monitored, queryable data.
Capture disclosures into consistent, versioned schemas.
Process high volumes of announcements automatically.
Industrialise repeatable extraction across document sets.
Structured Data Extraction turns unstructured documents into clean, repeatable datasets. Agents do the heavy lifting and call OCR only when it is needed. You define the schema and rules, and data is mapped to it reliably, with every action governed by approvals and a full audit trail.
Agents orchestrate the process and decide when to use OCR, table detection and other models. Results are validated and mapped to your schema automatically. When confidence is low, exceptions are raised for review, and humans can always override and provide final sign-off.
PDFs, Word, Markdown and common images (JPEG, PNG, GIF, BMP, TIFF, ICO). Images and Word files are converted to PDF for consistent processing, and mixed batches are handled together.
Use SFTP, S3/Blob/GCS, UI upload or REST API. Files are validated, de-duplicated and virus-scanned automatically, and you can configure watch folders, batching and retention rules.
It is hybrid by design. Confidence thresholds route uncertain items to a review queue where reviewers approve or correct them, and every change is logged with approver, time and reason code for full accountability.
SED is optimised for complex, real-world tables using structure-aware parsing and validations, with OCR when needed. Routine layouts achieve strong accuracy and tricky edge cases go to review, while feedback and rules improve results over time.
We'll show you how it works.