Conversation
Dropping descriptive phrases ("Associates", "Legislative Services", "Policy
Group", ...) groups variants of one firm's name, but when it leaves a bare
surname it merges different firms: three firms that share a surname and
differ only in those descriptors all normalized to the surname alone. Keep
the descriptor in that case. Also space out "&" (a name written "A&B" no
longer becomes "AAND B"), split glued "...Associates", and match phrases as
whole words.
On dev's 17,639 names this splits only those three firms (and one client
listed both with and without "Associates"); people's names are unaffected.
The TypeScript mirror produces identical output for all of them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sted Registration pages (Summary.aspx) were read only for their disclosure links and then discarded, so a lobbyist whose firm files their disclosures had no record at all. - Parse registration pages in full (lobbyists, employing firms, clients, dates, amounts; error pages return None) and store them in lobbyingRegistrations, one doc per registrant and year. Before 2019 the portal served several identical pages per firm per year and its lobbyist links pointed at those duplicates, so they collapse into one doc and lobbyist links are only kept from 2019. - Weekly and backfill runs save the registration from each page they fetch. - compute_stats lists every registered lobbyist and firm (with their firms or lobbyists, latest SoS registration link and per-year registrations for profile pages), and counts distinct registered individuals for the overview's lobbyist stat. - Each filing records the SoS disclosure page it came from (disclosureUrl). - migrate_lobbying_data.py gains a registrations phase and a phase that re-fetches archived pages that are really portal error pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Overview: "Individual lobbyists" now counts distinct registered individuals (previously lobbyist-year registrations), with an explainer noting that before 2019 firm-employed lobbyists registered under their firm. - Lobbyists list: includes lobbyists whose employer files for them, shows the entity they're registered under, and labels types as "Individual lobbyist" / "Lobbyist entity" (the SoS term for firms and organizations that employ lobbyists). - Profile page: entities a lobbyist is registered under, lobbyists an entity registered (linked when they have a page), a link to each year's SoS registration page, and a note in place of an empty bills table for lobbyists whose employer files for them. - Filing tables (profile, client, bill pages and the bill card) link each row to the SoS disclosure page it came from. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
…ration A transient network outage left pages unprocessed in the write phase (they keep stale data), and the error-page scan counted pages it couldn't download as fine. Retry each page's whole read/parse/write with backoff, and list pages the scan couldn't check instead of silently passing them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fake docs and normalization cases now use invented people and firms that keep the same patterns (commas, "&", typos, a shared surname), instead of real lobbyists and firms. Assertions on committed SoS fixture pages still match those pages; their comments are now generic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seedLobbyingStats.ts duplicated compute_stats in writer.py, which already runs after every scraper run that writes data. The copy had no tests, had fallen behind (no registrations), and read whole collections in one query, the pattern that timed out at this data size. Running it would have reverted the stats. For a manual recompute, `python3 scrape.py --mode stats` recomputes lobbyingMeta from existing data. The ingestion doc now describes it, the current name normalization steps, the registrations collection and the filing source URL. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nesanders
force-pushed
the
lobbying-registrations
branch
from
October 2, 2026 00:15
97d8d75 to
947a195
Compare
normalize.ts mirrored lobbying-scraper/normalize.py from when the scraper ran as a Cloud Function. Since the scraper moved to Python, every stored name is produced by normalize.py, and nothing imports the TypeScript copy (it was only re-exported from functions/src/lobbying/index.ts, which isn't loaded). Remove it and its test, and update the ingestion doc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From checking the registration-based pages on dev: - List each employing entity once, with its latest spelling, instead of once per spelling it used over the years. - On the profile of a lobbyist whose employer files for them, show only the note under Bills (no empty table, filing count or Clients heading). - Style the bill page's Source heading like the others, and render the SoS link arrow as text rather than an emoji. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Lists every registered lobbyist, including those whose firm files their disclosures and links lobbying data back to its source on the SoS website.
Summary.aspx) were only read for disclosure links; they're now parsed and saved tolobbyingRegistrations(lobbyists, employing entities, clients, dates). Before 2019 the portal served one identical page per firm per lobbyist, so those collapse into one doc.Removed: TypeScript copies of the Python pipeline. The scraper was first written as a Cloud Function, then moved to Python on Cloud Run because the SoS portal's firewall blocks Node's HTTP client. Two TypeScript files kept duplicating Python logic after that:
scripts/firebase-admin/seedLobbyingStats.ts— a manual script copyingcompute_stats(writer.py), which already runs after every scraper run that writes data. Every stats change had to be made twice; the copy had no tests, didn't include registrations (running it would have reverted this PR's stats), and read whole collections in one query, the pattern that timed out at current data size. For a manual recompute, usepython3 scrape.py --mode stats.functions/src/lobbying/normalize.ts— a copy ofnormalize.py. Nothing imported it (it was only re-exported fromfunctions/src/lobbying/index.ts, which isn't loaded); every stored name comes from the Python version.migrate_lobbying_data.pygainsregistrationsandrefetch-errorsphases (22 archived registration pages are portal error pages).Checklist
SosSourceLinkis a plain link)firestore.indexes.json(Please do not only create indexes through the Firebase Web UI, even though the error messages may reccommend it - indexes created this way may be obliterated by subsequent deploys) (N/A — the new single-doc read needs no index;lobbyingRegistrationsisn't read by the frontend)Screenshots
Rebuilt Lobbyist explorer page
Search across lobbyists and employers
Profile page for individual lobbyist filing as part of firm
Lobbyist page - note the additional SoS source links
Known issues
writeandregistrationsphases ofmigrate_lobbying_data.pyon prod, thenpython3 scrape.py --mode stats(no deletes needed; doc IDs don't change).Steps to test/reproduce
cd lobbying-scraper && python3 -m pytest.