• 2 Posts
  • 8 Comments
Joined 23 days ago
cake
Cake day: September 14th, 2026

help-circle

  • Why though? AI is already part of how a lot of engineers work, and people are already using it to game traditional coding interviews. so trying to enforce “pretend AI doesn’t exist” feels a bit like giving a kid a calculator every day, then deciding the exam should test whether they can hide from the calculator.

    i’d rather accept that the tool exists and test the part that still matters: can they reason, verify, catch bad output, understand the codebase, and make a safe change?

    the AI shouldn’t be the thing being tested. the engineer’s judgment should be.




  • yeah, i get that concern.

    i don’t think the point should be “can this person produce code fastest with an llm.” that would be pretty bleak.

    for me the interesting part is almost the opposite: can they still make good engineering decisions with AI in the loop? understand tradeoffs, reject bad suggestions, preserve the design, know when the generated fix is technically valid but wrong for the codebase.

    if the assessment only rewards output, then yeah, it just turns people into cogs.


  • i think we’re talking about slightly different things, i’m not suggesting the llm should decide whether someone is a good engineer or a good fit. i wouldn’t trust that either.

    the interesting part to me is putting someone in the environment they would actually work in, repo, ticket, tests, AI available and seeing how they reason through it.

    do they understand what the AI gives them? question it? verify it? know where to look when it’s wrong?

    the conversation afterwards can still be the most important part. the repo just gives you something concrete to have that conversation about.



  • this comment actually sent me down a rabbit hole lol

    the “score the trace, not just the diff” part is the bit i keep coming back to. the final fix is easy to measure, but the interesting signal is probably what they inspected, what they tried, whether they verified the AI suggestion, etc.

    i’ve been building something around this since that thread and that’s still the part i haven’t figured out properly - Groundwork but i’m curious, how much of the investigation do you capture before it starts becoming fake/surveillance-y?