This summer, I joined Text Blaze as a software engineering intern. It was my first time working in a real production codebase, and my first task was literally just a ticket: here’s a function with a bug, here’s what we want it to do, here’s what it’s doing instead.
Then I actually opened the codebase.
But actually seeing the function, seeing the file it’s in, and seeing a sea of many other files absolutely terrified me.
Instead of just working, I convinced myself I needed to understand everything first. What’s the current behavior? Who calls this function? When? Why? With what data? What calls those callers?
The thing is, this wasn’t even a complex fix. The final pull request ended up being a few extra lines and a test file. Had I approached this ticket with a better mindset, I would’ve finished it in half the time.
So this is the article I wish I read before touching production code. Maybe you’re starting your own internship, or want to make your first open-source contribution. Either way, I hope I can make your first pull request a little simpler and less stressful :)
Don’t assume infallibility.
The first thought I had in my head as I read this code was that this was the work of some senior engineer with 10 years of experience on me. If there was a bug in it, it was probably a REALLY nasty one. Surely not something simple!
And even though that’s probably true, production code isn’t sacred. It’s still written by people (or even an AI agent), and the bug is just that, a bug.
Don’t let the size or maturity of a codebase convince you the problem in front of you is more complicated than it looks.
You don’t need to master the entire codebase.
In fact, no one has.
If your scope is a single function, just start there.
You may need to trace it’s callers and inspect the data flowing into it, or understand another layer of abstraction, but you absolutely shouldn’t start there.
Reading more of the codebase will always give you a better understanding, but there’s a difference between a “better understanding,” and a “more relevant understanding.”
Test your assumptions.
Print the shape of data at every single step. Use a debugger. Inspect actual values.
It’s very easy to stare at code and try to speculate about what btrfs_fs_info looks like, but printing it will likely give you very immediate answers to whatever uncertainties you have.
Research language features as you go.
Take a look at this C++ code:
template <typename T>
concept Serializable = requires(T t) {
{ t.serialize() } -> std::convertible_to<std::string>;
};
template <Serializable T>
void save(const T& obj);
If you aren’t familiar with C++, this looks super cursed. In fact, even if you do know C++ well, this definitely takes some time to understand.
Seeing this, it might be very tempting to convince yourself you need to fully understand C++ templates before continuing. But usually that’s a trap. You’re better off just researching exactly what you’re looking at.
If you want to stay on-task, just ask an AI chatbot (if you want to), or go through the language’s documentation.
Write tests
One fear I had with my code was simply, what if I break everything else?
There are some obvious guardrails: git branches, code review, CI, and hopefully, an extensive existing test suite.
Run the tests, and if they’re still green (like they were before, hopefully), you’re most likely fine.
Of course, tests aren’t infallible. But existing tests simply tell you that you didn’t break anything else. It tells you nothing about whether or not you actually fixed the bug, or implemented the functionality.
That’s why you likely need to write your own tests too.
In my experience, it’s much easier to test as you go, instead one-shotting everything.
And when I say tests, I mean proper tests. Print statements are good for debugging, but production code needs real, reproducible unit tests. Don’t forget the edge cases too.
Just get started
Are you on another branch? Okay. Don’t overthink, just start printing things.