Overview
DataAnchor is a proposed data layer that sits between unstructured business information — documents, emails, support tickets, messages — and the AI applications, APIs and databases that consume it. Its job is to turn text into clean, validated, structured JSON while keeping sensitive identifiers out of downstream systems.
Today, DataAnchor exists as this website and a working browser-based prototype of the pipeline. The prototype is intentionally rule-based: it shows the shape of the product — the stages, the output contract, the validation behaviour — without depending on a model or a backend.
Prototype status
Pipeline
The planned product processes every input through six stages. The prototype implements the first five.
- Input — normalise line endings and whitespace, enforce limits.
- Detect — find identifiers and classify the document.
- Redact — replace identifiers with consistent, typed placeholders.
- Extract — map text to schema fields, normalising dates, amounts and enums.
- Validate — check types, formats and required fields; set
processedorneeds_review. - Route — deliver the record to its destination (planned).
In the planned service, extraction operates on redacted text so that no later stage sees raw identifiers. In the browser prototype, rules read the original text locally and only redacted values are written to the output.
Demo engine
Working in prototypeThe engine is about 1,500 lines of dependency-free TypeScript in src/lib/engine. It is deterministic — the same input and mode always yield the same output — and it performs no network I/O. It is covered by automated tests in tests/engine.test.ts.
Detectors
When two detections overlap, the higher-priority kind wins (secret, email, IBAN, card, SSN, IP, phone, name), so an email address is never partially redacted as a name.
Classification
Each matching rule adds its weight to one document type. The highest total wins if it reaches 3; otherwise the input is reported as unknown and validated against the generic schema with status needs_review.
Normalisation
- Dates — ISO dates, “October 30, 2026”, “30 Oct 2026”, DD.MM.YYYY and MM/DD/YYYY become
YYYY-MM-DD. Impossible dates are rejected. Dates without a year are kept as written and fail validation rather than receiving an invented year. Ambiguous slash dates are read as MM/DD/YYYY and noted. - Amounts — US (1,250.00) and European (1.250,00) formats become numbers. A labelled total is preferred, then an amount introduced by words such as “for” or “charged”, then the largest value.
- Currencies — explicit ISO 4217 codes are used as-is. Symbols are mapped (€ → EUR, £ → GBP); “$” is assumed to be USD and that assumption is reported.
- Enums — priority, ticket category, email intent and payment method come from keyword scoring and are only set when a keyword matches.
Schemas
Each document type has a versioned schema. Fields marked required must validate for the result to be processed; PII fields are redacted or masked in the output according to the selected mode.
invoice@v1
Invoices and payment requests, including remittance details.
support_ticket@v1
Bug reports, incidents and help-desk requests.
customer_email@v1
Inbound customer correspondence.
generic@v1
Fallback when no document type has enough matching signals.
Redaction modes
Placeholders are consistent within one document: repeated values share a number, and a first or last name on its own is linked to the full name detected earlier. Free-text fields such as subject keep their wording, with any identifiers inside them replaced.
Illustrative API
In designNo API or SDK has been published. The examples below describe the intended interface so that the output contract can be discussed; names, parameters and endpoints are placeholders.
from dataanchor import DataAnchor # placeholder package name client = DataAnchor(api_key="YOUR_API_KEY")result = client.extract(input=text, schema="invoice@v1", redact_pii=True) result.status # "processed" | "needs_review"result.data # dict matching invoice@v1result.redactions # list of {type, token}result.warnings # assumptions such as "$ assumed USD"Privacy model
- The demo runs entirely in the browser tab. Submitted text is never sent to a server.
- Nothing is written to cookies, local storage or IndexedDB; text disappears on reload.
- In production the site sends a Content-Security-Policy with
connect-src 'self', so the browser itself blocks scripts from contacting other origins. - The console measures requests to other origins during each run using the Resource Timing API and shows the count.
- No analytics, trackers or third-party scripts are loaded. Fonts are self-hosted.
Limitations
- English-language text only; three document types plus a generic fallback.
- Person names are found only in the contexts listed above. Names elsewhere are not redacted.
- Postal addresses, dates of birth and most national ID formats are not detected.
- Organisation names rely on labels, common suffixes (Inc., Labs, GmbH…) or phrases such as “from …”.
- Pasted text only — no PDFs, images or attachments.
- The prototype is a demonstration. It must not be used as a compliance control or relied on to remove all sensitive information.
See also the privacy notice and legal disclaimer below.