Everything Beady does starts with the pipelines that pull from 100,000+ public sources — company registries, court and enforcement records, sanctions and watchlists, news and social. You would own the part that keeps them flowing, complete and current.
What you would do
- Build and maintain collectors for registries, court systems, regulators, sanctions lists and media
- Keep sources alive when they change format, rate-limit, or simply go down
- Design storage and indexing for hundreds of millions of entities and their documents
- Make every record traceable back to the page it came from
What we look for
- Strong Python (or Go / Java) and SQL, comfortable with queues, schedulers and distributed workers
- Real experience with large-scale crawling or ETL and with messy, inconsistent data
- A bias for correctness — a record we miss is not a bug, it is a compliance failure
Send your CV and two lines about why this role to hello@beady.ai.