ai engineering / about me / software engineering
How I Code Review a 50,000-Line AI-Generated Feature
Austin Starks breaks down how he reviews a massive, AI-written pull request before it goes into NexusTrade, the platform running over $40,000 of his own money across two live brokerage accounts.
The example is real: a pull request that touched 419 files with over 50,000 lines added. Rather than review it as one block, he split it into 11 smaller pull requests, cut everything that was formatting churn rather than substance, and worked through each piece against the actual goal of the change -- tests, data model changes, and the deploy path -- before rolling it out behind a feature flag and watching production for issues.
He also shows the actual cost of a small mistake: a single misplaced minus sign in the order-submission code, from a real commit fixing a sign convention that could have sent trades to the broker with the wrong order type.
Figures on screen (lines of code, account values, user count) were pulled live from NexusTrade's own systems and GitHub the morning this was filmed.
Transcript
0:00I'm a senior AI engineer and I've vibe coded over 1.8 million lines of code.
0:07Not only that, but the software system I created is managing over $40,000 of my real actual money.
0:15I have users which depend on the software to manage their money as well and any bug, the smallest mistakes, even a negative sign, any bug can have catastrophic consequences.
0:30Despite that, I've been able to manage building and maintaining an extremely large software system, partially in fact due to my masters in software engineering from Carnegie Mellon, which according to U.S. News and World Report is literally the best school in the entire world for software engineering and the best school in the entire world for artificial intelligence.
0:55Here's how I code review a brand new feature like this, which is over 50,000 lines of code.
1:04Step one, I split up the PR into maintainable, reviewable parts.
1:09So one 50,000 line PR can get split up into 11 different PRs.
1:16Number two, I tried my best to reduce things like formatting chart and other irrelevant stuff, which might have made the PR artificially inflated.
1:29Number three, I orient myself on what exactly is this PR, what is the actual goal I'm trying to accomplish?
1:38Then I take a look at the code and I figure out if there are unrelated things according to the actual goal of the PR and I'll either remove them or move them to a separate PR, which they can be reviewed independently.
1:53I check out the tests and make sure we're testing the actual core business logic.
1:58I look at any artifacts like screenshots or other test results like golden files to make sure that we're actually testing the software appropriately.
2:07I review the data models, especially if we're adding or removing any fields and making sure that there is a plan so that when I deploy the software, it doesn't break in production, such as having reasonable defaults or coercing within the application.
2:22Once everything looks good, I review the PR and I do this across the stack, starting with the new backend functions and then updating existing backend functions and then finally doing the frontend.
2:37And ideally for a new feature, I'll also have it feature flagged so that after I deploy it, it's not immediately visible to the user.
2:48Finally, I roll it out and then I watch for any new system alerts, any new bugs or issues, and I also check my analytics to make sure users are actually interacting with the new surface.
3:05By doing all these things, I ship fast, reduce production downtime, and respond to issues as soon as they happen.
3:15Follow-up for more.
Join the conversation
Loading conversation…