Prompt

How do I set up a SEC filings data API for universe construction with normalized identifiers and auditability?

Finance / Financial Data2 observationsLast seen Sep 29, 2026

Latest observation

Sep 29, 2026OpenAI APIWeb search: off

To set up a SEC filings data API for universe construction with normalized identifiers and auditability, design it as a small data platform rather than just a file downloader.

1) Define the core goals

Your API should support:

  • Universe construction: “Which companies belong in my investable universe on date X?”
  • Normalized identifiers: map SEC-specific IDs to persistent master IDs.
  • Auditability: every universe membership should be explainable from source filings and rules used.
  • Point-in-time correctness: avoid look-ahead bias.

2) Use a canonical entity model

SEC data is messy because identifiers change and one company may have many filings and identifiers.

Recommended canonical IDs

Use a layered approach:

  • entity_id: your internal stable ID for the economic entity.
  • cik: SEC Central Index Key.
  • ticker: optional, time-varying.
  • figi / permno / LEI: optional external mappings if available.
  • adsh: accession number for a specific filing.
  • filed_at and period_end: for point-in-time logic.

Entity normalization rules

Create an entity resolution service that:

  • ingests SEC CIK/company name/filing data,
  • tracks name changes, mergers, delistings,
  • maintains a history table of identifier mappings.

Example structure:

  • entities
  • entity_identifiers
  • entity_name_history
  • entity_events (merge, spin-off, delist, reincorporation)

This lets you answer:

  • “What did this company map to on 2022-03-15?”
  • “Which current entity corresponds to CIK 0000320193?”

3) Separate raw, normalized, and derived layers

Use a medallion-style architecture:

Raw layer

Store original SEC artifacts exactly as received:

  • filing metadata
  • XBRL instance documents
  • filing HTML
  • filing facts
  • submissions JSON
  • company facts JSON

Keep:

  • source URL
  • download timestamp
  • hash/checksum
  • response headers if possible

This is critical for auditability.

Normalized layer

Transform raw filings into structured tables:

  • filings
  • filing_facts
  • financial_statements
  • company_profile
  • entity_mappings

Normalize:

  • units,
  • periods,
  • accession numbers,
  • CIK formatting,
  • fiscal calendars,
  • currencies.

Derived layer

Build universe-ready outputs:

  • universe_membership
  • eligibility_flags
  • screening_metrics
  • issuer_coverage

4) Build an auditable universe construction pipeline

Universe construction should be rule-based and versioned.

Example workflow

  1. Select all entities with a filing history.
  2. Apply eligibility rules:
    • exchange listing status
    • domicile
    • sector
    • minimum market cap
    • reporting completeness
    • filing recency
  3. Apply exclusions:
    • bankrupt/delisted
    • OTC-only
    • foreign private issuer, if excluded
  4. Freeze results as-of a date.

Store the decision trail

For each membership decision, persist:

  • entity_id
  • as_of_date
  • rule_set_version
  • included boolean
  • reason_codes
  • source_filing_ids
  • computed_at
  • data_snapshot_id

This gives you full reproducibility.


5) Design API endpoints around point-in-time use

A useful API should expose both raw data and universe outputs.

Suggested endpoints

Entity lookup

  • GET /entities/{entity_id}
  • GET /entities?cik=...
  • GET /entities?name=...&as_of=...

Identifier mapping

  • GET /identifiers/{cik}
  • GET /entity-identifiers?as_of=...

Filings

  • GET /filings?entity_id=...&from=...&to=...
  • GET /filings/{adsh}

Filing facts

  • GET /filings/{adsh}/facts
  • GET /facts?entity_id=...&concept=...&as_of=...

Universe

  • GET /universes/{universe_id}/members?as_of=...
  • GET /universes/{universe_id}/rules
  • GET /universes/{universe_id}/audit/{entity_id}?as_of=...

Snapshots

  • GET /snapshots/{snapshot_id}
  • GET /snapshots/{snapshot_id}/entities

6) Make identifier normalization explicit

Do not bury mapping logic inside downstream analytics.

Example identifier mapping table

entity_idciktickerstart_dateend_datesource
E1230000320193AAPL1980-12-12nullSEC submissions
E4560001018724TSLA2010-06-29nullSEC submissions

Track time ranges so you can do:

  • current mapping,
  • historical mapping,
  • point-in-time mapping.

Important

Tickers are not stable identifiers. Use them only as display attributes or temporary lookup keys.


7) Include lineage and reproducibility metadata

Every response used in research should be traceable.

Store:

  • source system
  • source file/path
  • ingest time
  • transformation version
  • rule version
  • checksum/hash
  • parent snapshot

Example lineage chain

SEC submission JSON → raw object → normalized filing record → universe membership record

This is what makes your pipeline auditable.


8) Use a snapshot strategy for research consistency

For backtests, don’t read live data directly.

Snapshot types

  • daily snapshot of filings and mappings
  • monthly universe snapshot
  • event-driven snapshot after ingestion cycles

Each snapshot should be immutable.

This avoids changes in:

  • amended filings,
  • late corrections,
  • mapping updates,
  • ticker changes.

9) Handle amendments and restatements carefully

SEC filings can be amended (10-K/A, 10-Q/A).

Best practice

  • Keep all filings.
  • Mark amendment relationships.
  • Choose a policy:
    • latest filing wins,
    • original filing only,
    • amended filing supersedes after filing date.

Store:

  • is_amendment
  • amends_adsh
  • supersedes_adsh
  • effective_from
  • effective_to

For universe construction, make the policy explicit and versioned.


10) Recommended data model

A practical minimal schema:

entities

  • entity_id PK
  • primary_cik
  • legal_name
  • country
  • created_at
  • updated_at

entity_identifiers

  • entity_id
  • id_type (cik, ticker, lei, etc.)
  • id_value
  • start_date
  • end_date
  • source
  • confidence

filings

  • adsh PK
  • entity_id
  • form_type
  • filed_at
  • period_end
  • accession
  • accepted_at
  • source_url
  • file_hash

filing_facts

  • adsh
  • concept
  • value
  • unit
  • frame
  • start_date
  • end_date
  • confidence

universe_rulesets

  • ruleset_id
  • name
  • version
  • definition_json
  • created_at

universe_membership

  • universe_id
  • entity_id
  • as_of_date
  • included
  • reason_codes
  • ruleset_id
  • snapshot_id
  • computed_at

11) Add quality controls

You’ll want automated checks such as:

  • duplicate CIKs
  • duplicate accession numbers
  • invalid date ranges
  • missing entity mappings
  • ticker overlaps
  • filing date after snapshot date
  • inconsistent currencies/units
  • orphaned filings

Log data quality metrics per ingest run.


12) Operational stack suggestion

A common setup:

  • Ingestion: Python + scheduled jobs / Airflow / Dagster
  • Storage: object storage for raw docs, Postgres for metadata
  • Analytics: DuckDB / Spark / warehouse for derived datasets
  • API: FastAPI
  • Metadata/lineage: custom tables or OpenLineage-compatible tooling
  • Caching: Redis or CDN for hot endpoints

13) Security and compliance

SEC data is public, but your platform may still need:

  • API authentication for internal use
  • rate limiting
  • audit logs for all data mutations
  • role-based access to snapshot publishing
  • immutable raw storage permissions

14) Example universe construction output

For an as-of date, your API response should look like this:

{
  "universe_id": "us_large_cap_equities",
  "as_of": "2024-12-31",
  "ruleset_id": "ruleset_v3.2",
  "snapshot_id": "snap_20241231_01",
  "members": [
    {
      "entity_id": "E123",
      "cik": "0000320193",
      "ticker": "AAPL",
      "included": true,
      "reason_codes": ["listed", "us_domiciled", "market_cap_pass"],
      "source_filings": ["0000320193-24-000010"]
    }
  ]
}

15) Best-practice principles

If you only remember a few things:

  1. Keep raw SEC data immutable
  2. Use stable internal entity IDs
  3. Track identifier history with effective dates
  4. Version all rules and snapshots
  5. Record why every entity was included or excluded
  6. Use point-in-time data only

If you want, I can also provide:

  • a reference database schema,
  • a FastAPI endpoint design,
  • or a step-by-step ingestion pipeline for SEC submissions and company facts.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (2 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.