g3ddy
AI Strategy8 min · 11 Oct 2026

The essence of 513 screenshots: what actually matters in AI-era engineering

513 screenshots distilled into one read: ten high-leverage ideas, a minimal toolkit, three builds and a 30-day plan for AI-era engineers.

Geddy
Geddy
Senior Web Engineer / Lead
0:00
Lots of screenshots stacked with spacing in 3d perspective as an explode style preview of a stacked screenshots in one queue with one highlighted and pulled aside that shows terminal window with glowing typing cursor. Futuristic, abstract background.

AI-generated content. This article was generated by AI from material I collected and has not been fully reviewed by me. It may contain errors — verify anything you rely on. More in the disclaimer.

TL;DR30 sec
  • The harness — context, tools, verification, retries — decides whether an AI feature ships; the model is the easy part
  • Spec first, test first, one feature at a time: vibe coding produces demos, specs produce products
  • Run 3–5 agents in parallel on git worktrees and never let an agent grade its own work — a second session catches what the first missed
  • Turn every repeated correction into a CLAUDE.md rule and every repeated prompt into a committed skill or command
  • Autonomy is earned, not granted: 20 logged runs at 95% verified pass rate before anything runs unattended, and nothing ships, posts or pays without you
  • Sell deployed systems and documented setups, not demos — price the outcome, skip the bidding wars, and route cheap work to cheap models Agents write the code; your job is now the system that verifies it.

Between January and September 2026 I saved 513 screenshots of threads, repos, charts and job posts about AI-era engineering. I never opened most of them a second time.

This is what's left after sorting them into nine long articles and then keeping only the parts that change how you work, build or sell.

The 10 highest-leverage ideas

  1. The harness matters more than the model. Most AI pilots that die in production never had one: a system that gathers context, calls tools, verifies and retries. Do this: before tweaking prompts, check what the model sees, what tools it has, and what verifies its output.
  2. Write a spec before the agent writes code. Vibe coding gets you a demo. A constitution → specify → plan → tasks → implement flow gets you a product. Do this: for any real feature, use Plan Mode or Spec Kit before a single line is generated.
  3. Verification is what separates engineers now, not code volume. Agents guess at unclear requirements unless a failing test anchors them. Do this: have the agent write a failing test first, then make it pass. One feature at a time.
  4. Run agents in parallel. The Claude Code team calls 3–5 sessions on git worktrees its single biggest productivity gain. Do this: git worktree add one tree per task, one Claude session in each.
  5. CLAUDE.md is memory that learns from corrections. A rule added right after a correction measurably lowers repeat mistakes. Do this: after each correction, say "update CLAUDE.md so you don't make that mistake again". Keep the file under a page.
  6. A prompt you retype every week should be a skill or command. Commands are just markdown files with $ARGUMENTS. Do this: name them as verbs (/grill-me, /triage-issue), commit them to git, and read anyone else's before you install it.
  7. No agent should grade its own work. A second session or a second model catches a different class of mistakes. Do this: before every PR, run "Grill me on these changes — no PR until I pass your test" in a separate session.
  8. Hooks do deterministic guardrails at zero inference cost. They run even when you forget to ask. Do this: add PreToolUse hooks that block rm, prune, push and dangerous git commands, and run lint and typecheck after edits.
  9. Treat cost as a routing problem. Bulk work goes to cheap models, hard reasoning to frontier ones, and everything in the context gets compressed. Do this: track token spend per workflow and re-price it against open-weight models every quarter.
  10. Autonomy has to be earned with evidence. A task type runs unattended only after 20 logged runs at a 95% verified pass rate. Do this: agents produce drafts and PRs. Nothing gets sent, paid, posted or merged without you.

The shift in one screen

Area Before Now
Engineering Author writing every line Orchestrator running parallel agents behind specs, tests and review gates
Code review "Claude did it" You own and can defend every line of an AI-assisted PR
Security A gate after the fact Inside the loop: pattern catch on edit, diff review each turn, validation at commit
Hiring bar Can talk about RAG and agents Has shipped a system that holds up under real load and fails gracefully
Roles Frontend / backend Forward Deployed Engineer, AI Product Engineer, Agent Engineer, Operations CTO
Senior careers One permanent role Fractional, interim or advisory work chosen on purpose for leverage
Business Adopt AI because of FOMO Fix the process first, because "AI does not fix broken processes. It makes them run faster."
Search SEO only SEO plus GEO: getting cited by AI answer engines
Solo operator One person does everything One person directs and reviews while agents do the recurring legwork on a schedule

The minimal toolkit

Name What Why it's worth it
Superpowers Skill set that enforces brainstorm → plan → TDD → review Discipline with zero config, in every session
Agent Skills (official) Anthropic's skills repo, including frontend-design The reference for writing a SKILL.md, and it stops UI from looking templated
Karpathy-inspired CLAUDE.md One-file CLAUDE.md: think first, keep it simple, make surgical changes The best starting memory file if you don't have one
claude-code-best-practice Starter system of agents, commands, memory, hooks and skills A production-ready template you can fork
Awesome Claude Code Directory of skills, plugins, hooks and tools Look here before you build, because someone has probably solved it
claude-mem Persistent memory across sessions Stops you re-explaining the project every session
headroom Compresses tool output, logs and RAG chunks 60–95% fewer tokens. Runs locally as a library, proxy or MCP server
Deep Agents MIT harness that copies Claude Code's architecture for any model The same design, fully transparent. Worth studying
MCP Standard way to connect models to tools and data Gives the agent live access to GitHub, databases, Figma and browsers instead of copy-paste
Ollama Local model runner that speaks the Anthropic Messages API A fallback when you hit caps or work offline, via ANTHROPIC_BASE_URL
shadcn/ui Component primitives Claude assembles from them instead of inventing UI
Taste Skill + Hallmark Two rule sets against "AI slop" Bans the purple-gradient, centred-hero, Inter-everywhere look
Playwright Browser automation and screenshots Closes the design loop: screenshot, compare to the reference, iterate
pgvector Vectors inside Postgres Keeps RAG next to your app data, with SQL filters and row-level security
n8n Self-hostable workflow engine Adds webhooks, retries and scheduling around headless agents
geo-seo-claude Open-source GEO audit Checks AI-search visibility, and you can sell it as a service

Three builds worth doing this month

1. Chief-of-staff morning brief — Claude Code headless + Gmail/Calendar/GitHub MCP + cron.

  • An agents/chief-of-staff/ repo containing SOUL.md, memory/ and a brief skill with a fixed template. Credentials stay in a .envrc and are read-only.
  • Run it by hand with claude -p "/brief" --max-turns 20 until it's useful, then schedule it. Log every correction in FEEDBACK-LOG.md.

2. Droplet ops watchdog — cron + claude -p + a shell allowlist.

  • A collector script gathers docker ps, docker logs --since 1h, disk, memory, TLS certificate expiry and HTTP checks. A triage skill returns OK, WARN or FAIL with a suggested command.
  • It stays quiet on OK. It runs as non-root, a hook blocks rm/prune/down/push, and after two weeks of correct diagnoses it may take one safe action.

3. "Chat with your docs" in Next.js — LiteParse → pgvector hybrid → reranker → cited answers.

  • Chunk by headings and store doc_id, page, bbox, section. Retrieve the hybrid top-50, merge with reciprocal-rank fusion, then rerank down to 5–8.
  • Require [doc:page] citations and an explicit "not found". Write 30 golden questions before you tune anything.

Selling it: positioning essentials

Sell production, not demos. Package one deployed feature as a "what broke, how I found it, what I did next" story. That's exactly where CTOs say candidates fail.

Getting hired or getting clients is a four-part system: mindset and consistency, positioning, selling yourself, and technical skill. Most senior engineers only work on the last one.

Headline formula: [Role] who [does X with AI] for [outcome]. Describe your practice with the Capable / Adoptive / Transformative vocabulary.

Your setup is the proof. A documented Claude Code setup (CLAUDE.md, skills, hooks), Lighthouse numbers and a model cost benchmark show judgment, not just claims.

Price the outcome, not the hours. Follow-ups within 24 hours. Invoices chased. Releases shipped. Don't sell "time saved" — that's only capacity.

Skip the bidding wars. Upwork gets 50+ AI-drafted proposals per job within 30 minutes. Reach out directly and compete on proof and specialisation.

Formulas & numbers worth remembering

  • CAC = (marketing + sales) / number of new customers. Compare it with lifetime value before you scale a channel.
  • AI ROI = (hours saved × fully-loaded hourly cost) ÷ AI spend. The same stack gave 6.45x in SF ($180/hr), 3.2x in Austin ($90/hr) and 1.0x offshore ($28/hr).
  • GEO: only 11% of domains get cited by both ChatGPT and Google AI Overviews. AI-referred traffic is up +527% YoY and converts 4.4x higher than organic.
  • GEO signals: brand mentions correlate 3x more strongly with AI citation than backlinks do. Gartner projects a 50% drop in traditional search traffic by 2028.
  • Subsidised plans: ChatGPT Pro ($200/mo) allows up to ~$14,000 of usage (70x), and Claude Max 20x up to ~$8,000 (40x). The subsidy "will end".
  • Open-weight pricing: DeepSeek V4 Pro reaches 79% of Claude Opus 4.8's benchmark score at about 5% of the price ($186 vs $3,700). GLM-5 claims ~97% of Opus 4.5 at ~14% of the price.
Geddy
Geddy
Senior Web Engineer / Lead

Engineering leadership • AI innovation • Product thinking. 20+ years of web engineering, from independent contractor to engineering leader. Passionate about developer experience and product engineering.

Follow on LinkedIn
Keep reading