AI coding assistants

Match the task to how closely you check the code

An AI coding assistant is only as safe as the checking you attach to it, and the right amount of checking depends on the task. Below is the matrix we use, a 30-minute worksheet to test any assistant on your own repository, and five failure checks that each come with their own fix.

Published by ZeroTwo. The matrix is our rule of thumb, not a study.

How to read each cell

Accept
Take it and move on.
Skim
Read the diff for plausibility.
Read
Read every line.
Test
Run a test before you merge.
Review
A second person signs off.
Prototype
May be a throwaway. Never production.
Reject
Write this one by hand.

The task and trust matrix

Read across as the cost of being wrong rises. Read down as the tool does more on its own.

Verification level for each coding mode, from completion to app builder, across four tasks, from a throwaway script to a security or PII path.
Mode, and how much it doesThrowaway scriptLibrary codeProduction hotfixSecurity or PII path
Completiona line at a timeAcceptAccept freelySkimSkim the diffReadRead every lineReviewPair with a reviewer
Chat in the editora function at a timeAcceptAccept freelyReadRead it and lintTestRequire a testRejectWrite it by hand
Repository agentseveral files, run by the toolTestRun it and smoke testTestRun the full suiteReviewSign-off and reviewRejectSkip the agent
App buildera whole appPrototypePrototype onlyRejectDon’tRejectDon’tRejectDon’t

Modes, not brands

The mode is how you use a tool for one task. It is not a property of the product: Claude Code’s own page lists the terminal, IDE extensions, the web and Slack, so one name can sit in several rows. We therefore sort by mode and name no tool as belonging to exactly one. For which editing surface fits your constraints, see our guide to AI code editors.

  • Completion

    A line or a small block

    What you verify
    Plausibility. You read each suggestion before accepting it.
    Use it for
    Keystroke savings, boilerplate, parameter lists and repeated patterns.
    Skip it when
    Architecture choices, anything you cannot read in two seconds, security-sensitive paths.
  • Chat in the editor

    A function or a block

    What you verify
    The code and its tests. You ask, read, and edit.
    Use it for
    Explaining unfamiliar code, scaffolding a function, rewriting a block, translating between languages.
    Skip it when
    Multi-file refactors and system-level reasoning.
  • Repository agent

    Several files, with the tool running commands

    What you verify
    A finished diff and the test run. You review a result, not a typing session.
    Use it for
    Routine refactors, dependency upgrades, generating tests, a small ticket with a clear definition of done.
    Skip it when
    Security-sensitive paths, and anything where being wrong costs more than doing it by hand.
  • App builder

    A whole working prototype

    What you verify
    That it is a prototype. Everything else is still unreviewed.
    Use it for
    Demos, internal tools, one-pagers and weekend experiments.
    Skip it when
    Systems with real users. A demo that looks live is still a prototype.

Assistant or agent: what changes for you

With an assistant you steer step by step and read as you go. With an agent you hand over a task and review a finished diff. The work you do changes from reading suggestions to reviewing a result.

A repository agent suits a task with a clear, testable definition of done, a repository whose tests run on every change, and a small blast radius if it is wrong. Read about when a repository agent is the right next step.

One task, three modes (an illustration, not a test)

“Add rate limiting to this endpoint”

  1. Step 1

    Completion

    Predicts the next few lines as you write the handler. You still write the structure.

    You read each suggestion as it appears.

  2. Step 2

    Chat in the editor

    Rewrites the handler when asked, edits the route, and updates the test in the same panel.

    You read the diff and run the test.

  3. Step 3

    Repository agent

    Adds the middleware, registers it in config, writes a test, runs the suite and posts a diff.

    You review a finished diff, and the cost of a miss now spreads across files.

The agent path is less typing, but the verification has to follow the autonomy you granted. That is the whole idea of the matrix.

Do assistants make developers faster?

It depends on the task and on how well you already know the code. The four figures below are labelled by what kind of evidence they are, because that is what people drop when they quote them.

Observed in a controlled experiment
55.8%faster
95 professional developers, all familiar with JavaScript, timed writing an HTTP server. The group with GitHub Copilot finished 55.8% faster. A small, fresh, well-specified task.
Read at arXiv and GitHub
Observed in a randomised trial, 2025
19%longer
16 experienced open-source developers, 246 real issues in repositories they knew, AI randomly allowed or forbidden. METR now marks these results out of date.
Read at METR
Reported in a survey, 2025
84%use or plan to use AI tools
Of respondents to Stack Overflow’s developer survey, 84% use or plan to use AI tools and 51% of professional developers use them daily.
Read at Stack Overflow
Reported in the same survey
46%distrust
46% of developers said they distrust the accuracy of AI tools and 33% said they trust it. About 3% said they highly trust the output.
Read at Stack Overflow

The 2025 slowdown, and what METR said next

METR’s developers expected AI to make them 24% faster and, after being slowed down, still believed it had made them 20% faster. In February 2026 METR wrote that its follow-up data is an unreliable signal, because developers who did not want to work without AI were dropping out, and that it believes developers are more sped up now than in early 2025. The update is worth reading in full.

What follows for you

No published figure is about your repository, your conventions or the tool you are considering. The one reliable number is the one you measure, and the people in the 2025 trial show that how fast it feels is not that number. The worksheet below is how to get it.

Evaluate an assistant on your own code in 30 minutes

Three steps, one rubric with anchors so two people score the same diff alike, and a decision rule. Everything below runs on tasks from your repository.

  1. 5 min

    Pick four tasks from your own repository

    A refactor of about 50 lines, a bug you already fixed and can grade against, a small new endpoint or flag, and a test file for an existing function. Tests are the cheapest place to find out whether the assistant has read your code.

  2. 4 x 5 min

    Run each in the mode you would really use

    Note the mode, then after each task answer the five failure checks below. Do not average the tasks together yet; a tool can be fine at tests and poor at refactors.

  3. 5 min

    Score each task, then decide per task type

    Score the four dimensions from 1 to 5 using the anchors below, then apply the decision rule. The answer is per mode and per task type, not one verdict on the tool.

Score each task from 1 to 5

Scoring anchors for four dimensions, giving the meaning of a score of 1, 3 and 5. Use 2 and 4 for in between.
DimensionScore 1Score 3Score 5
Diff qualityDoes not run, or breaks existing tests.Runs, but needs fixes before the tests pass.Passes your tests with no edits.
Context fitIgnores your conventions, or invents files and APIs.Mostly follows conventions, with a few mismatches.Matches your naming, structure and existing helpers.
Time and costTook more of your time than writing it by hand.About break-even on time and spend.Clearly less time and spend than by hand.
ReviewabilityParts of the diff are unexplained or indefensible.Mostly defensible, with a few lines to rework.You could defend every line in code review.

Decision rule (our defaults, set your own bar). If context fit or reviewability scores below 3 on any task, do not use that mode for that type of task. If every score is 4 or above, keep it. Anything in between: re-run with a narrower task or a lower-autonomy mode.

A completed sample, with invented numbers

A fictional assistant, run in repository-agent mode on a fictional repository. The scores are made up to show how the rule and the checks work together. They say nothing about any real product.

Invented sample worksheet: four tasks with scores for diff quality, context fit, time and cost and reviewability, the failed check and its fix, and the outcome computed from the decision rule.
TaskDiffContextTime and costReviewFailed check and fixOutcome
Refactor: extract a validation helper (about 50 lines)5444NoneKeep this mode for this task type
Bug: a closed issue about duplicate webhook events3243Check 2: kept code it could not explainAsked for a line-by-line explanation, still could not explain one block, wrote that part by handDo not use this mode for this task type
Feature: a new read-only report endpoint4434Check 3: skipped a testWrote the test before merging; it caught an off-by-oneRe-run with a narrower task or a lower-autonomy mode
Tests: unit tests for an existing parser4545NoneKeep this mode for this task type

After each task: five failure checks

Answer each with yes or no. A yes is always a failure, and each failure has its own fix. There is no count that triggers a rollback. Log which check failed and in which mode: the same check failing again and again on one task type is your signal to change mode.

Five failure checks. For each: the statement to answer yes or no, what a yes means, and the fix for that failure.
Check (a yes is a failure)What a yes meansFix for that failure
1. The diff took me longer to review than it would have taken to write.The mode is too autonomous for this task, or the ask was too broad.Split the ask into smaller pieces, or move this task type one row up the matrix, toward more manual work.
2. I kept code I cannot explain in plain English.You are about to own code you do not understand.Ask the assistant to explain it line by line. Delete any line you still cannot explain and write that part yourself.
3. I skipped a test I would normally have written.The speed came from skipping the check that would catch the assistant’s mistakes.Write the test now, before you merge, and run it against the diff.
4. I committed the first answer without a second pass.You have no comparison, so you cannot tell a good answer from a plausible one.Ask for an alternative or a review of the diff, and compare the two before committing.
5. I would not send this diff to review without reading it again.You do not yet trust your own read of it.Read it again now, top to bottom, before you open the pull request.

The seat price is not the whole cost

Attention is a cost

An assistant that needs a full read of every diff costs you attention even when the seat is cheap, and one you trust on routine work costs less of it. The worksheet’s time and cost score is where that shows up. We list no vendor prices here because they change; check each vendor’s own pricing page, and ask what a free tier’s quota counts and what happens when it runs out.

When code cannot leave your network

The constraint then is where inference runs. Ask whether the tool can point at a model endpoint you control, what it sends and stores, and whether your policy allows each model. Local and self-hosted setups trade vendor convenience for control, and you take on running the inference yourself. Run the worksheet on that setup too, because quality varies with the model behind it.

Where ZeroTwo fits

ZeroTwo is a chat and agent workspace you keep next to your editor assistant, with access to 60+ models. Use it for the prompt that needs a different model than your editor ships with. It does not complete code inside your files. See the ZeroTwo code agent, or every plan and allowance.

  • Free

    100 free credits to start

    $0/moNo card required

  • Plus

    Access to code agent

    $14.99/mo$11.99/mo billed annually

  • Pro

    Access to ZeroCode

    $29.99/mo$26.99/mo billed annually

Questions about AI coding assistants

What is an AI coding assistant?

Software that uses language models to read, write, refactor and review code, from line-level completion to agents that work through a whole task. The useful distinction is not the product but the mode you use it in for a given task, because each mode asks for a different amount of checking.

How much should I trust the code it writes?

As much as your checking justifies. In Stack Overflow’s 2025 survey, 46% of developers said they distrust the accuracy of AI tools and 33% said they trust it, and only about 3% highly trust the output. Use the matrix on this page: the higher the cost of being wrong, the more of the diff you read, test and have reviewed.

Do AI coding assistants make developers faster?

It depends on the task, and the evidence is mixed and dated. In a controlled experiment published in 2022 and 2023, developers using GitHub Copilot completed a small JavaScript server task 55.8% faster. In METR’s 2025 trial, experienced open-source developers took 19% longer on real issues in their own repositories with AI allowed, and METR has since said its newer data is an unreliable signal. The way to know for your code is to time yourself with the worksheet on this page.

How do I evaluate an AI coding assistant fairly?

Run four real tasks from your own repository, a refactor, a closed bug, a small feature and a test file, in the mode you would actually use. After each, answer the five failure checks, then score diff quality, context fit, time and cost, and reviewability from 1 to 5. Decide per mode and per task type rather than giving the tool one verdict.

What is the difference between an AI coding assistant and an AI coding agent?

With an assistant you steer step by step and read each suggestion or edit. With an agent you hand over a task, it reads the repository, edits files and runs commands, and you review the finished diff. They are modes, not product names: a single product can be used either way.

Are AI coding assistants free?

Many vendors offer a free tier and some tools are open source with bring-your-own-model, but limits and prices change, so check each vendor’s current page. Ask what the quota counts and what happens when it runs out. ZeroTwo’s Free plan starts with 100 free credits plus 20 bonus credits each day you log in, in a chat workspace; it is not editor completion.

Where does ZeroTwo fit if I already use an editor assistant?

ZeroTwo is a chat and agent workspace you keep next to your editor, with access to 60+ models. A code agent is included from the Plus plan and ZeroCode from Pro. It does not complete code inside your files.

Keep reading

  • AI code editors

    Which editing surface fits your constraints: an extension, an AI-first editor, or a task agent beside your editor.

  • AI coding agent

    ZeroTwo’s page on handing a repository task to an agent.

  • Best AI for coding

    Our broader guide to choosing an AI tool for coding.

Run the worksheet, then ask the hard prompts in ZeroTwo.