SSR Codex
An annotation tool that lets senior engineers teach SSR's AI assistant what the documents never said. Live and rolling out to an internal team.
My role
Lead designer, By the Pixel
Team
Me, a PM, and a lead engineer
Deliverables
Annotation dashboard, Annotation editor, Guided entry flow


About Codex
Teaching Smitty what the documents never said
SSR's engineers answer technical questions with help from an internal AI assistant called Smitty. It searches the firm's library of mechanical, electrical, and plumbing standards, calculators, and design guides, then answers from what it finds. The catch: those documents were written for experienced engineers. They leave out context a newer engineer needs, and some standards were never written down at all, so Smitty comes up short exactly where help matters most.
Codex is the tool senior engineers use to fix that. They work through a document, find what's missing, and add it as annotations Smitty can draw on. As annotations accumulate and link related documents together, Codex starts to work less like a pile of notes and more like a map of what the firm actually knows.
I led the product design, working with a PM and a lead engineer at By the Pixel: the annotation dashboard, the annotation editor, the metadata model, and the flow an engineer follows to create an entry.
The problem
The knowledge lived in heads, not on pages
Ask Smitty about panel boards and it can retrieve the panel board standard and summarize it. But if the standard never spelled out that toilet rooms require GFI outlets, Smitty can't tell you that. The document never said it. That knowledge lived in a senior engineer's head, not on the page.
Working through documents this way surfaced a second problem: some standards had no home document at all. The firm was carrying knowledge it had never written down, and nobody was tracking the gap.
There are really two jobs in here. One is retrieval, finding the right document. The other is interpretation, understanding that document once you have it. The team chose to focus on interpretation first. User feedback pointed to context as the bigger gap, and it was the easier of the two to isolate and measure. Codex is the interpretation half.

Who uses it
Built for the firm's busiest people
The people doing this work are discipline leads, the experts who already carry the knowledge. They are also the firm's scarcest, busiest people, and that constraint shaped every decision that follows. If the tool cost them time, they would abandon it, and the whole effort would fail.
Codex has three roles. Super Admins manage the system, Technical Leaders own and edit content in their discipline, and Reviewers can weigh in without full editing rights. The permissions differ, but the approval process between them stays deliberately light. A heavy review chain would punish exactly the people the tool depends on.

The extraction conversation
A conversation, not a form
The obvious design here, handing experts a form and asking them to write down what they know, fails with exactly these users. Writing context from a blank page is slow, and the firm's busiest people would do it once, badly, and never come back.
The better rhythm showed up in an observation session. A senior engineer uploaded a real document to Claude and worked through it section by section, building a companion document of the context the original left out. It worked, and it was slow: it took a lot of turns, the companion document had no connection to the source at the granular level where the context belonged, and chat was a clunky place to build a contextual layer. Underneath that, though, the rhythm was right. The AI proposed roughly fifty additions one at a time, and the engineer kept the source open and checked every suggestion against it.
I designed the interface around that rhythm and against those two failures. A conversation keeps the engineer in a reviewing posture rather than a data-entry one: they react to concrete proposals about a document they know cold, which is faster and more accurate than writing context from a blank page. It also keeps them honest. Because each step references the source, nobody is tempted to rubber-stamp whatever the AI produced. And the editor below is what fixes the rest: every annotation is anchored to the passage it belongs to, and the conversation lives beside the document instead of replacing it.
The AI does the legwork of finding what's missing. It never produces final numbers, and it never gets the last word: a person who knows the document reviews everything, and that person stays responsible for what's published.

The annotation editor
Comments in the margin, context in the thread
Fifty proposed annotations are only worth having if reviewing them costs less than writing them would have, so the editor is built to make reacting cheap. Starting a document is one click and an optional nudge. From the dashboard, an engineer hits Begin Annotating and can hand the LLM instructions before it reads: what to focus on, what this document tends to assume. Then the model reads the ingested document and drafts annotations on its own judgment.
The drafts land where the work is. Each annotation sits in the right margin beside the highlighted text it references, the way comments work in a Google Doc, and drafts are unmistakably yellow. The engineer works through them one at a time, and every annotation carries its own text input, so feedback happens directly next to the thing being addressed: accept the draft as written, or tell the model what's wrong and let it regenerate. Manual editing exists, but SSR wanted the LLM doing the writing, so the interface leads with feedback rather than a blank field. Approve an annotation and it turns green. The left sidebar holds the document's outline with a count of remaining drafts per section, so how much is left is always one glance away.
Engineers aren't limited to reacting, either. Highlight any passage and add an annotation manually; the person chooses where, and the LLM still takes the first stab at what.
Each annotation also carries its own working metadata. An optional expiration date marks information with a shelf life, so a value that will change gets revisited instead of quietly going stale. An annotation can be assigned to a colleague, for review or because the topic is really theirs. Links connect an annotation to related documents and Codex entries, the first threads of the knowledge map. And every annotation has a unique URL, so engineers can share one in an email or a chat without asking anyone to go dig for it.
Alongside the margins runs a global chat sidebar holding the whole conversation with the model, and individual annotations surface in that thread too. Context builds in one place instead of fragmenting across fifty small exchanges, so what the model learns in section 3 is still in the room by section 9.

The annotation dashboard
Counts and color, not progress bars
With dozens of documents in flight across disciplines, the dangerous state is not knowing what's done. The dashboard answers it: documents flow in from the firm's SharePoint and appear as un-started, ready for anyone to pick up or assign out. From there the dashboard shows every document's annotations, who owns them, what state they're in, and whether they've been published. A "My Documents" tab shows an engineer only what's assigned to them, and separate filters cut the full list by other criteria, like a discipline's backlog.
One decision here mattered more than it looks: how to show where a document stands. Each row carries a fraction, approved annotations over total, next to a colored dot that stays yellow until every annotation is approved and turns green at 100%. The denominator is real, the annotations that actually exist, not a guess at a finish line, so the number never lies the way a progress bar sitting at 80% would.
Approval isn't the whole story, though. An annotation can be approved but not yet published, and until it's published, Smitty can't see it. The mental model is a lot like git: approved-but-unpublished is committed-but-unmerged, real work that isn't live yet. So the dashboard tracks publishing separately from approval: an Unpublished badge flags documents carrying approved work that hasn't shipped, and a last-published date column shows when each document's context actually reached Smitty. Between the dot, the badge, and the date, an engineer can tell at a glance which documents still need work.
Outcome
From a pile of notes to a map
Codex is live and rolling out to a small internal team at SSR. The first release centers on the dashboard and the annotation flow, with linking and the richer knowledge-graph features scoped for later phases. The standard it was designed to: a senior engineer works through what the AI proposes against a document they know cold, publishes context Smitty can use, and never signs off on anything they didn't check.
What I'd do differently
The extraction conversation is fifty questions long because that's what the observation session produced, and I designed for that length rather than questioning it. The open question is whether the AI should read the whole document before asking anything, and lead with the ten gaps it's least sure about. Fewer, better questions would cost the firm's busiest people less, and the rollout is where we'll find out.
More Work




Vora
Privacy-first agentic AI.


Scarlet>Fire
Stream the Grateful Dead archive.
Contact
Fill out the form and I’ll get back to you ASAP. Or email me directly at jesse.bisignano@gmail.com.










