I turn messy websites and other unstructured sources into data agents can answer from.
Most of what a company knows sits in web pages and documents that no model can use as they are. I build the pipelines that turn them into data an AI assistant can answer from, with every answer linked to its source.
Start here
Live demo. I crawled a real company's public website and put an assistant on top of it. Ask it anything; every answer cites the page it came from.
ask.hellon.rio (going live soon)
What I've built since 2023
- A crawler that only re-processes pages that changed, so keeping a large site current costs close to nothing.
- Boilerplate removal at the site level: sibling pages share a template, so I detect it and strip it deterministically, before any model runs.
- Knowledge extraction where each fact keeps a verified link to the passage it came from.
- A RAG assistant with a bounded agent loop. Every change to what enters the context has to win an eval against the previous version.
Before this, ten years of web products in JavaScript, React and Node, and a physics degree.
Work with me
Open to contract or full-time work, remote, with small teams building agents or data-extraction products.
LinkedIn · GitHub