Market-data lake

lake.political_filings dataset schema

Column types and research semantics for NexusTrade’s lake.political_filings dataset. Read the supported schema before submitting a bounded SQL query.

Query name
lake.political_filings
Grain
year
Columns
26

Columns and types

These columns come from the same versioned logical catalog used by the query validator and engine. They describe the fields you can query. Some historical rows have missing values; fields absent from older shards can return typed NULL.

chamberVARCHAR
docIdVARCHAR
filerFirstVARCHAR
filerLastVARCHAR
filerSuffixVARCHAR
stateDistrictVARCHAR
filingDateDATE
availableAtTIMESTAMP
availabilitySourceVARCHAR
sourceUrlVARCHAR
rawArchiveKeyVARCHAR
rawSha256VARCHAR
parseMethodVARCHAR
extractionStatusVARCHAR
failureReasonVARCHAR
extractedRowsINTEGER
extractionModelVARCHAR
contractVersionVARCHAR
ocrArchiveKeyVARCHAR
amendedReportDateDATE
reportDateDATE
processedAtTIMESTAMP
filerKeyVARCHAR
memberIdVARCHAR
displayNameVARCHAR
identitySourceVARCHAR

Current logical schema for lake.political_filings

Coverage and timing

The logical view is backed by canonical manifest-resolved data. Query this name rather than guessing a physical database collection or object-storage path.

House and Senate periodic transaction reports, one row per indexed filing. `availableAt` is the point-in-time column: a filing may inform a decision only when `availableAt <= as_of`. It comes from the date named in `availabilitySource`, the House disclosure index filing date or the Senate eFD filed date. Never use a transaction date as availability. Filings whose trades were not extracted are rows too, so completeness is a query. `extractionStatus` is `ok` when the filing's trades are in `political_trades`, `failed` with a `failureReason` when extraction did not finish (the daily job retries it), and `unsupported` for a source form with no gated extraction. `ok` means a filing was read, not that every value is right: the rows are what a model read from the filed page, and known-hard cells are dense micro-print dates and checkbox columns that share a bound. `contractVersion` and `extractionModel` say what read a filing, so rows read by an older contract can be found; a filing is read again and its rows replaced whole when an operator names it. Filer identity: `memberId` is the member's Bioguide ID from congress-legislators and `displayName` is their official name. One person has one `memberId` across every spelling of the filed name and across both chambers, so identify, group and filter politicians by `memberId` and show `displayName`; `filerFirst` and `filerLast` are the name as filed and vary between filings. `identitySource` is `legislators` or `override` for a member, `non_member` for a House employee or candidate who never served (`memberId` NULL), and `unresolved` when no member could be decided. When the question names a person, a `# Politician identity` section carries the resolved `memberId` — filter by that literal and never re-derive identity from `displayName`. `parseMethod` is how the filing was read: `text` and `scan` are House PDFs with and without a usable text layer, `html` and `paper` are Senate electronic and paper reports. `amendedReportDate` is set on a filing that amends an earlier report. `reportDate` is the date a Senate eFD report title prints; a report and its numbered amendments share it, and it can differ from `filingDate`.

Run a limited query

Use a registered API key with lake scope. Select needed columns, bind user-supplied values and restrict dates when the table has a time dimension. Query results are durable parts with a schema and manifest rather than an unbounded in-memory array.

This reference documents the schema and its meaning. Run queries in your signed-in workspace; this page does not execute SQL or show private results and datasets.

sql
SELECT "chamber", "docId", "filerFirst"
FROM lake.political_filings
LIMIT 20;