Application cockpit
A local tool for a senior job search. It drafts a tailored CV and cover letter, checks both against my real history, and never submits anything. Built to find out what actually keeps an AI draft honest.
Client
WeDoUX
Service
AI Product Development
Date
June - July 2026

Project Overview
Context
Senior roles I'd actually take appear a few times a month. Tailoring each application meant pasting my background, my writing voice and the job description into a chat that had forgotten all three, then checking the output line by line for things I never did. The application cockpit keeps that context resident. It sources roles from seven applicant-tracking systems and drafts the CV as structured data and the letter as prose. Both are verified before a designed PDF is rendered. It never submits: there is no code path that could. Built solo, in a public repo.
The real problem
The brief I gave myself was better cover letters. The letters were only the symptom. My writing rules went verbatim into every prompt, and they name "it's not X, it's Y" as a construction to kill. Every letter the tool had produced used it: sixteen letters, sixteen violations, all sixteen over the 300-word ceiling at a median of 538. The CV path had gates. The letter path had nothing that measured anything, so no change could be shown to help. The system was full of controls that weren't controlling anything, and it had no instrument to notice.
My role
All of it: the product, the pipeline, the gates, the evaluation scripts and the experiments. Claude Code wrote the implementation; I designed each piece and reviewed every change before it landed. Five of the scanner's provider handlers began as code Greg Nudelman gave me during the course, and the repo credits them as his. No users and no client.
System-level decisions
Every guarantee is code that either runs or doesn't. A language gate blocks generation when a posting asks for a language I haven't declared, and quotes the requirement back. A numeric gate traces every number in a draft to my CV, the posting or my profile, because team sizes are where interview liability lives and a model invents them politely.
My writing rules became code too: banned constructions as patterns, word counts, one implementation imported by both the live path and the offline tests. Two copies drift, and then a draft's fate depends on which door it came through.
Every experiment was pre-registered. Before each run I wrote down what would justify shipping, including a human read of the letters with veto power over the metrics.
Outcome
The change that mattered most was to a measuring instrument. A check that a letter delivers the figures its plan committed to reported one commitment dropped in two runs of five. Both letters had delivered it, written out as "eighteen designers" where the check looked for "18". The check had fired corrective retries at correct letters, degrading the prose it was there to protect. It now reads "eighteen" as a number. Three other safeguards failed the same quiet way, including a CI workflow that had never run once.
All three experiments stopped on their own criteria. The two-step letter architecture made every commitment land, five runs out of five, and flattened the prose: sentence-length variance of 10.81 against a bar of 11.61, set before the first API call. It ships switched off.
The limit: the tool records every draft and every gate result, and nothing about whether an application got a reply. Everything here measures against the standard I set. None of it is validated against what works.


Key Highlights
16 / 16: letters over the word ceiling before any fix
4: safeguards found not working, each by checking the control itself
3 / 3: pre-registered experiments stopped by their own criteria
578: automated tests, all passing


Go Back

