PrivyusPrivyus

Technical Architecture

The hard part of Privyus is the data work under the screens. Records about one person sit in many sources, and each one must join a single node with a link to the original record. The demo shows layer 4 (the product experience) with fixed and illustrative data. The real product adds three layers under it: collect, resolve, and store.

Data
Public records
Federal sources first, then the 50 states
Technology
Available today
Standard databases, connectors, and LLM extraction
First version
One topic, federal sources
A goal for a small senior team

Data Volume

Federal data alone runs to hundreds of millions of rows, and new filings arrive every day.

~108

Hundreds of millions of rows

Itemized campaign contributions (FEC)

~107

Tens of millions of rows

Member votes, each member on each roll call

~107

Tens of millions of pages

Filings, statements, and the Congressional Record, as text and embeddings

→109

Approaching billions

Edges after resolution: every contribution, vote, contact, trip, and meeting as a link

×n

Several times all of the above

All 50 state legislatures added

Also in the federal setMillions of rows · Lobbying filings and their contacts, issues, and clientsMillions of rows · Bills, actions, and cosponsor records

Order-of-magnitude estimates for federal data only, over several election cycles. Exact counts depend on the years and sources covered.

Hard Problems at This Scale

Name Matching Without IDs

One person appears under a different name in each source. Most records carry no common identifier.

A Source on Every Edge

Each edge keeps its source document with the fetch time, the URL, and a content hash.

Sub-Second Answers

Queries stay fast while new filings arrive every day and the graph changes under them.

Four System Layers

Each layer feeds the next one, and each layer can grow on its own. Records move from the public sources on the left to the app on the right. The pipeline keeps three tiers of tables, and only the serving tier is visible to the app.
Data flowHuman step

Hover or select a layer to see its detail

Public sourcesOriginals, with fetch time, URL, and content hashcongress.gov APIAPIRoll call votesXMLSenate LDAAPIFARABulkOpenFECAPIGift and travelPDFFinancial disclosuresPDFStatementsHTMLGovInfoBulkLayer 1CollectAPI connectorscongress.gov, LDA, FECBulk file loadersVotes, FARA, GovInfoScrapersHTML, PDF, and OCROrchestratorSchedules and retriesLayer 2Resolve and ConnectNormalizeNames, dates, amountsEntity resolutionPublic IDs, fuzzy, LLMRelationship extractionTyped edges with sourcesReview queueUncertain matchesLayer 3StoreRaw document storeEvery original filingGraph databaseNodes and edgesSearch and vector indexFull text, embeddingsAnalytics warehouseCounts and trendsApp databaseUsers, watchlistsLayer 4This demoQuery and DisplayAI query agentTools: graph, searchResponseAnswer, delta, citationsWeb appDashboard and graphWorkspaceCheckpoints and alertsrawAs delivered, never editedbuildRebuilt by the pipelineservingRead-only, after checksThe app readsthe serving tier only

Layer 2

Resolve and Connect

Records about the same person become one node, and each record becomes typed edges. This step is what a customer pays for because no single public database does it.

Normalize
Clean names, dates, amounts, and addresses into one format.
Entity Resolution
Start from public IDs (Bioguide, FEC, LDA). Use fuzzy matching and an LLM check for records without IDs. Send uncertain matches to a human review queue. Never guess silently.
Relationship Extraction
Each record becomes typed edges: member voted on bill, lobbyist contacted office, committee contributed to member, member traveled to place. Each edge keeps its source document and its date.
Social Network Precedent
Large social networks resolve identities the same way. They combine many weak signals into one confident identity.

Layer 2 · Resolve and connect

Entity Resolution

Raw records name the same person in many ways. Privyus must turn them into one node. It starts from public IDs, uses fuzzy matching and an LLM check where no ID exists, and sends uncertain matches to a person. It never guesses silently. Large social networks solve the same problem when they combine many weak signals into one confident identity.

Sen. Ellen Hartley in five sources

Decided by public IDDecided by modelDecided by reviewer
  • congress.gov

    Hartley, Ellen

    1.00Public ID
  • OpenFEC

    HARTLEY, ELLEN M.

    1.00Public ID
  • Senate roll call

    Hartley (R-OH)

    1.00Public ID
  • Senate LDA

    HARTLEY, ELLEN

    0.93Model
  • Travel filing

    Sen. E. Hartley

    0.97Reviewer
EH

Sen. Ellen Hartley

R-OH · person

privyus_id per_01H…

One node with one stable ID

ID CrosswalkOne row for each known identifier
crosswalk (key_type, key_value, privyus_id, confidence, decided_by)  bioguide        H001234            per_01H…   1.00  public_id  fec_candidate   S6OH00123          per_01H…   1.00  public_id  senate_lis      S401               per_01H…   1.00  public_id  lda_contact     "HARTLEY, ELLEN"   per_01H…   0.93  model  travel_filer    "Sen. E. Hartley"  per_01H…   0.97  reviewer

Stable IDs

A rebuild reuses an existing ID. It creates a new ID only for a new entity, so links, watchlists, and saved checkpoints never break.

Human Review Queue

Uncertain matches go to a human reviewer. Reviewer decisions (“same person”, “not the same person”) live in Postgres, and every rebuild applies them first.

Release Checks

The release compares the new IDs with the last release. A sudden jump in merges or new IDs stops the release.

Layer 3 · Store

Serving Data Model

The product is the connections, so the serving tier stores nodes and edges, not one wide row per person.
entitiesperson | org | bill | event | place
PKidtext
typeenum
nametext
subtitletext
partytext
statetext
photo_urltext
edgesevery link, with its source and date
FKsrc_idtext
FKdst_idtext
typeenum
event_datedate
amountnumeric
FKsource_doc_idtext
confidencereal
Sorted by (src_id, type, event_date). A second copy sorted by (dst_id, type).
documentsthe source chip on every answer
PKdoc_idtext
sourcetext
urltext
fetched_attimestamp
content_hashtext
titletext
entity_stats“Votes 42”, “Trips 6”, “Meetings 11”
FKentity_idtext
categorytext
countint
votesthe seat chart and the map
vote_idtext
FKmember_idtext
positionenum
  • A meeting or a trip is a node (type = event). The attendees connect to it. This is the Meetings → Private meeting → attendees chain in the demo.
  • Every edge has a source_doc_id. This is the source chip on every answer.
  • Money from a company PAC or its employees links to the company with an edge that has a confidence score and a source. It is never a silent merge.

Layer 3 · Store

One-Hop Queries

Each click in the demo maps to a query that reads one sorted range. A columnar database (for example, ClickHouse) or Postgres answers each one in milliseconds. Open a link to see the same click in the demo, next to its query.
Each demo action and the one-hop query that answers it
Demo actionQuerySee it in the demo
Select a bill, see its supportersedges WHERE dst_id = bill AND type IN (sponsored, cosponsored) JOIN entitiesS. 456 supporters(opens the demo in a new tab)
Select a member, see the categoriesentity_stats WHERE entity_id = memberSen. Ellen Hartley(opens the demo in a new tab)
Open Meetingsedges WHERE src_id = member AND type = attended JOIN entities -- eventsMeetings(opens the demo in a new tab)
Open a private meetingedges WHERE dst_id = meeting AND type = attendedPrivate meeting(opens the demo in a new tab)
“Which defense contractors gave to this senator?”edges WHERE dst_id = member AND type = contributed JOIN entities -- orgsDefense contractors(opens the demo in a new tab)
Open a sourcedocuments WHERE doc_id = edge.source_doc_idVisitor log source(opens the demo in a new tab)

Layer 4 · Query and display

How the Graph Is Drawn

The agent returns data in one fixed shape, and the app draws it. The agent never writes free-form queries.
  1. Template

    The agent maps the question to a tested query template and fills in the parameters. It does not write free-form SQL.

  2. Graph Delta

    It returns the nodes and edges to add, with the answer text, the citations, and the steps. The demo already uses this exact shape.

  3. Layout

    The app keeps the existing nodes in place and lays out the new nodes around the selected node with a force layout.

  4. Map Arcs

    The map uses the same edges. Each organization and member has a state, so a contribution becomes an arc from one state to another.

Agent ResponsePlay it in the demo
// trimmed: 3 of 5 meetings, 4 of 5 steps{  "id": "hartley-meetings",  "steps": [    { "label": "Focus Meetings",      "delta": { "add": { "entities": [], "edges": [] },                 "focus": "cat-meetings" } },    { "label": "Query meeting records",      "delta": { "add": { "entities": [], "edges": [] } } },    { "label": "Populate meetings",      "delta": { "add": {        "entities": ["meeting-private", "meeting-meridian",                     "meeting-committee"],        "edges": ["cat-meetings-meeting-private",                  "cat-meetings-meeting-meridian",                  "cat-meetings-meeting-committee"] } } },    { "label": "Render",      "delta": { "add": { "entities": [], "edges": [] },                 "focus": "cat-meetings",                 "expand": "cat-meetings" } }  ],  "answer": "Five meetings appear in Hartley’s office records, …",  "citations": ["disc-meetings", "disc-private"]}
Referenced Recordsentities · edges · documents
// entities{ "id": "meeting-private", "type": "detail",  "label": "Private meeting", "sublabel": "Jun 17, 2026" }// edges{ "id": "cat-meetings-meeting-private",  "source": "cat-meetings", "target": "meeting-private",  "label": "meeting", "date": "2026-06-17",  "sourceIds": ["disc-private"] }// documents{ "id": "disc-private", "kind": "DISCLOSURE",  "title": "Visitor log · Jun 17, 2026",  "geo": { "from": "DC", "to": "OH" } }

One Shared Shape

Each step completes when its part of the graph appears. The edge carries the id of its source document, so the answer can show a source chip. The document carries the two states that the map draws as an arc.

A real backend returns this same shape, so it can replace the fixed data file without a rebuild of the screens.

Layer 4 · Query and display

Agent Access

Planned
Privyus can serve AI agents from the same core as the app, through a REST API and an MCP server. Every result carries its sources, so an agent can cite the filing behind each claim. Agents will do a growing share of the research that analysts do today.
Request flowPlanned

Hover or select a door to see its route

ClientsDoorsShared coreServing tierAnalystAsks in the web appCustomer softwareInternal tools and dashboardsAI agentsAny MCP clientFund research agentNewsroom agentCompliance agentGeneral AI assistantWeb appDashboard and graphThis demoREST APIJSON over HTTPSPlannedMCP serverSix tools for agentsPlannedOne core for every doorQuery templates + citationsTested query templatesNo free-form queriesOne answer shapeText, graph delta, citationsSource on every resultURL, fetch time, hashCoverage limitsThe same rules as the appServing tablesChecked before releaseentitiesedgesdocumentsentity_statsControlsOn every doorAPI keysOne per customerRate limitsPer key and per toolUsage logsCalls per keyAudit logsEach request and result
  • This demo

    Web App

    Analysts ask in plain words and explore the graph. Each answer shows its source chips.

  • Planned

    REST API

    Customer software calls the same query templates. Each response is JSON with the answer, the graph delta, and the source ids.

  • Planned

    MCP Server

    MCP (Model Context Protocol) is an open standard that lets AI assistants call outside tools. Any agent that supports it can search Privyus and cite each filing it uses. Each agent platform that connects brings Privyus records into the tools its users already use.

Agent Research Session

A fund’s research agent asks one question and makes five tool calls. Each result adds nodes to the same graph that the app draws. The records, dates, and amounts come from the demo data.

Agent Session5 tool calls · 2 sources

Question to the agent

Which defense contractors met with Sen. Ellen Hartley’s office or gave to her campaign in 2026?

  1. search_entities({ query: "Ellen Hartley" })
    per_hartleySen. Ellen HartleyR-OH
    1 person
  2. expand_connections({ id: "per_hartley", type: "attended", from: "2026-01-01" })
    evt_meridian_0224Meridian policy briefingFeb 24
    evt_ukraine_0429Ukraine embassy delegationApr 29
    evt_private_0617Private meetingJun 17
    3 events
  3. expand_connections({ id: "evt_private_0617", type: "attended_by" })
    per_vossClara VossMeridian Public Affairs
    per_pierceNathaniel PierceAegis Systems
    per_senDr. Priya SenAtlantic Security Institute
    3 people
  4. expand_connections({ id: "per_hartley", type: "contributed", from: "2026-01-01" })
    org_aegisAegis PAC$18,750 · Apr 8
    org_redwoodRedwood PAC$7,625 · Aug 21
    2 contributions
  5. get_sources({ edges: ["edg_private_pierce", "edg_aegis_hartley"] })
    [1] doc_disc_privateVisitor log · Jun 17, 2026
    urlhttps://…/visitor-log-0617.pdf
    fetched_at2026-06-18T09:14Z
    hashsha256:9f2c…41ab
    [2] doc_fec_hartleyFEC Schedule A · Aug 2026
    urlhttps://…/schedule-a/S6OH00123
    fetched_at2026-08-22T06:02Z
    hashsha256:5d07…c3e2
    2 documents

Answer

Nathaniel Pierce of Aegis Systems attended a private meeting in Hartley’s office on Jun 17, 2026 1. FEC receipts list 2026 contributions to her campaign from Aegis PAC ($18,750) and Redwood PAC ($7,625) 2.

[1] Visitor log · Jun 17, 2026[2] FEC Schedule A · Aug 2026
Graph Built by the AgentSame records in the demo
Waiting for the first tool result

Each tool result adds its nodes. The orange path leads to the record behind citation 1.

MCP Tools

The planned MCP tools, what each returns, and where the session above uses it
ToolWhat it returnsIn the session
search_entitiesPeople, organizations, bills, and filings that match a name or a topicCall 1
get_profileOne entity with its category counts (votes, trips, meetings, contributions)
expand_connectionsThe one-hop neighbors of an entity, filtered by edge type and dateCalls 2 to 4
get_sourcesThe original filings behind an edge, with URL, fetch time, and hashCall 5
find_pathThe shortest documented chain between two entities
watchNotifications when a new record about an entity arrives

Layer 3 · Store

Stores and Releases

Each store answers a different type of question. A release reaches the app only after its checks pass, and it can be rolled back.
The stores, what each holds, and the questions each answers
StoreHoldsAnswersExample technology
Raw document storeEvery original filing“Show me the source”Amazon S3
Graph databasePeople, organizations, bills, events, and edges“How is A connected to B?”Neo4j or Amazon Neptune
Search and vector indexFull text and embeddings“Find statements about Ukraine aid”OpenSearch or pgvector
Analytics warehouseCounts and trends over timeDashboard counters, Radar, trendsClickHouse, Snowflake, or BigQuery
App databaseUsers, watchlists, checkpointsThe personal workspacePostgres

Start with one Postgres database, and split out stores as the data grows. A columnar database or Postgres answers the one-hop queries above; a graph database becomes necessary only for multi-hop path questions.

Release Safety

  1. Parallel Build

    Each release builds new serving tables next to the live ones.

  2. Release Checks

    Row counts, ID churn, empty fields, and a fixed set of benchmark queries compared with the last release.

  3. Atomic Swap

    The new tables replace the live ones in one atomic swap.

  4. 30-Day Rollback

    The old tables stay for 30 days, so a rollback takes one command.

Incremental Updates

Each source also runs an incremental update, daily or hourly, that keeps its position with a cursor. This feeds the live activity feed and the watchlist alerts.

Risks and Mitigations

Each risk below is a known engineering or process problem, with a known way to reduce it.
Risks, why each matters, and how to reduce it
RiskWhy it mattersHow to reduce it
Entity resolution errorsA wrong link between two people is a false claim.Public IDs first, a confidence score on each match, human review for uncertain matches, and a source on every edge.
Inferred claims about real peopleDefamation and reputational risk.Show records, not conclusions. The product says “the visitor log lists X”, never “X influenced Y”. Legal review of the wording.
Scanned and messy documentsOCR and extraction errors.Schema checks, confidence scores, and a link to the original page.
Source changesA site changes its layout, and a scraper breaks.Monitoring per connector, alerts, and API sources first.
Gaps in the public recordPrivate meetings are often not disclosed.Be clear about coverage. Show what each source covers and does not cover.
LLM cost and speedEach question calls the model several times.Cache common queries, use templates, and use smaller models for extraction.
Data termsSome sites limit automated access.Check each source’s terms. Prefer official APIs and bulk files.

Phased Plan

Phase 1 builds the data foundation that every later phase needs. The analysis layers come last, because they need the history that phases 1 to 3 collect.
Phased plan: scope, team, and time for each phase
PhaseScopeTeamTime
Phase 0DemoThe current interactive demo, with fixed data.DoneDone
Phase 1Data foundationFederal sources: congress.gov, votes, LDA, FEC. Entity resolution for members of Congress and committees. Raw store and one Postgres database.2 to 3 engineersAbout 3 months
Phase 2First productThe AI query agent with citations, the graph and dashboard on real data, one tracked topic, watchlists and checkpoints, sign-on. Pilot with 3 to 5 design partners.3 to 4 engineers, 1 designerAbout 3 more months
Phase 3CoverageFARA, travel, financial disclosures, statements. The human review queue. Alerts.4 to 6 peopleOngoing
Phase 4Analysis layersSentiment over time, then prediction. Needs the history from phases 1 to 3.Adds data scienceAfter phase 3

These are planning estimates for a small senior team. They are not quotes.

Running cost in phases 1 and 2 is mostly cloud hosting and LLM usage. It grows with users and sources. A design partner pilot can run on a modest monthly budget.