Some lessons I've gathered along the way:
The blessed "Have you tried turning it off and on?" and its companion "just wait it out".
When you're immature, you get frustrated when a complicated ticket comes your way and you wish for that same ticket you can do with your eyes closed, exactly as the previous 200 times... But instead, you have to deal with some annoying customer/complicated network topology/ancient technology. This statement includes both the technical complexity as well as dealing with difficult people. It's instinctual to try to avoid these situations. The problem with this approach is that you'll learn a lot less doing only what you know. As usual, progress is located outside your comfort zone.
This is to be expected in some cases because bigger customers simply have more locations, equipment and links, so there are higher chances of something breaking. But here I'm talking about those annoying customers whose "Internet is always slow" ('cause they can't differentiate between a shit application/website and a problem with the Internet link), whose "settings are always wrong" ('cause they change their network every three weeks), and whose "email is always malfunctioning" ('cause they can't configure their accounts and passwords properly). It's hilarious when they threaten by saying they'll take their business to our competition. Yeah, I bet they can't wait to have you on board. That would actually both benefit us because we don't have to deal with you anymore, and impede the competition because they now have to. Win-win.
People get anchored too easily. In general, refrain from giving time estimates for when it will be fixed, especially in complicated cases.
Sure, this issue is resolved. But another one will turn up soon enough.
Big companies usually have their dedicated network people or even networking departments, which makes them the easiest to work with because they know the procedures, how things work, how things don't work, and most importantly, how things go from not working to working. Medium-sized companies are workable because they have at least one IT person to whom you can transfer the education of the user on how to connect to the correct WiFi, how to change the email settings, or how to forward a call. With small companies, everything stated before falls on you. Or the alternative: explaining to them that you're not their IT and that they should hire one, which they're always delighted to hear.
Sometimes laypeople complain just to see if something can be improved, but have no concrete evidence of what exactly doesn't work: "the internet is kinda sluggish", "the call quality is mediocre", "I'm sensing the WiFi signal has gotten worse". If you take these complaints too seriously, you end up chasing ghosts and troubleshooting something that works just fine for the level of service they're paying for - which is usually the most basic grade.
KissCo has a useful model of outlining levels of understanding a particular technology in their exam topics - describe, explain, configure, troubleshoot. You can try skipping steps, but you'll end up with a lot of joyful hours of pain - you can ask me how I know... Because I'll answer you: I got my first job in a troubleshooting role, so inevitably I ended up troubleshooting something I couldn't describe/explain, let alone configure. Aaaaah, the pleasure...
I first remember hearing about something similar from Russ White: "just because you can configure it doesn't mean you understand it". This one addresses the unavoidable trap of certifications (I myself also fell into): you want to shortcut it (for good and bad reasons) so you glance over the fundamentals and deeply understanding the material and jump to the syntax and configs which are more represented on the exam; because they easier to test and vendor certs are to some extent advertising brochures.
The other reason is that if you don't understand how it works, how it breaks and what it's designed for; you certainly don't understand how it affects the rest of the network - both working as it should and breaking as it shouldn't, but most certainly will. Sometimes it's about what you don't configure.
There's a reason network architect/designer positions are among the most senior roles. Young people (speaking from my friend's experience...) are captivated by sexy job titles like "architect", "red teaming", "cybersecurity"... These are not entry-level roles. "I want to get my ENWLSD so I can design networks." Yeah, sure. The ironic part is that it's regarded as the easiest CCNP Enterprise concentration cert. Operating, breaking, troubleshooting and fixing a lot of networks provide design insights.
In general, you'll be the one with more knowledge in this particular domain, so you're in charge of the conversation. People like to be taken care of. They've come to you, you're there to do this professionally, they're the customer. You don't go to your mechanic and tell him how to fix your car. Depending on the customer, on the other side can be an IT/network person with whom you'll be more on an equal level and troubleshooting efforts will be mutual. However, it shouldn't happen that the other side is more competent than you at your job.
During the troubleshoot, be wary of the other side trying to help, and as a result, leading you down the wrong trail. This one's especially important if you lack experience.
This allows the network to converge and machines (and people) to realize that something has changed. If you call the customer immediately, then you waste time waiting for them to check if they can browse the web, if the POS system is working, have the phones updated their status bar... If you wait a bit, by the time you contact them, they will have already figured out the network is fixed.
It's a real hassle to get someone who isn't sure which OS they're using ("the regular one I thinkkk?") to provide you with their IP or MAC address (please don't ask me how I know). If they're connected by wire, then almost always their distance to the router/switch doesn't matter, so you shouldn't care anyway. But you care if they're connected by WiFi and you're responsible for it. Instead of torturing yourself by explaining to them how to get to the CMD prompt and then spelling ipconfig, the more elegant solution is asking them to disconnect and reconnect again, which they're usually capable of. You can then determine where they are using logs, client uptime, or which device disappeared briefly on the connected devices list.
Some random error when installing a VM, software image impossible to find, installed version of the program doesn't have the feature you want...
Labs are useful to test solutions in a safe environment (All network engineers have test enviornments. Some even have production.). But some things aren't labbable due to size, complexity, cost... Eventually you have to push to production.
Labs are useful for learning. I enjoy labbing and understanding technologies from the ground up. But you learn the most on the job. For one, because some things aren't labbable (at least not at full scale); and secondly, there you'll encounter a mix of diverse technologies from different vendors with various perspectives, and when you add imaginative customers to that mix, you get something that never would have occurred to you on your own.
The network is a collection of devices linked together meant to communicate. Anthropomorphizing the network by calling it your "baby" develops an emotional attachment which in turn makes it harder to modify it upon realization that your "baby" is complicated, ugly and deformed. Plus, it's cringe (and not in a good way).
I occasionally come across comments about high-quality courses, books or resources that say things like "too dry for me" or "I learn better through videos." Even if the idea of learning styles were true (which is debatable), that's a poor attitude to have if your goal is to learn or improve your craft. The aim of these courses isn't to entertain you, but to teach you. So sink your teeth in - you'll come out stronger and smarter.
Medium-sized companies usually have an IT person or two who isn't a network person but is in charge of the network. The company "invested" in some SDN solution with a shiny dashboard. The complaint goes as follows: "Our dashboard is red and sometimes orange for this location, plz fix". When you ask them what the dashboard measures and how it examines the network status (latency, bandwidth, humidity of the server room, wind velocity), they usually respond: "red, sometimes orange... but all the other locations are green or yellow, plz fix".
In a (strictly) troubleshooting role, you have the privilege of bearing the blame for every bad interaction the customer has had in the past with your company (hey, who doesn't adore their ISP?). Maybe during service activation their questions weren't properly answered, the service isn't working exactly as they imagined, they're being overcharged for subpar service, or they'd already had the pleasure of interacting with the customer service IVR and then waiting on the phone for 10 minutes for an agent who'll listen to them half-assedly for 2 minutes... And eventually they land on you: the Internet doesn't work, the phone doesn't ring, the WiFi is terrible, etc. Rest assured they'll absolutely seize this chance to let you know about everything that pissed them off.
"My superiors in Germany/US/Hogwarts will hear about your lousy company!", "I'm 10 times smarter than every single one of you!", "I masturbate every Tuesday evening with your CEO!"... Emotionally incompetent people express their frustrations inappropriately. This partially speaks to their intelligence as well, because it's real smart to insult someone in charge of solving your problem. I straight up just ignore these ramblings and steer the conversation towards the troubleshoot - "how can we get this router working?", "when can you check the LEDs?", "who can the technician contact after arriving at the location?"...
You should listen to what they're saying and be empathetic. Here comes BUT. I've seen online some customer service advice that's just plain bullshit that translates to "be an agreeable punching bag, let them vent, listen to everything they have to say...". No, fuck off. I have other customers who'll be grateful for my effort. Criticism and suggestions are always welcome (or at least tolerable when it won't change anything), but lashing out, not gonna happen. I'm not even sure the big company compartmentalization, bureaucracy, inertness (and "revolving door" in this particular department) are even solvable, like by the laws of physics. Soooo, when they start their monologue and you know where they're going with it because you've heard it too many times before, save both of you some time by interrupting and keep the conversation troubleshoot centered: "There certainly have been mistakes made by us, I understand the frustrations, but how can we solve this concrete case?" Occasionally however, they're insufferable. Me, personally, I take pride in being the last person to hang up the phone. Or you can hang up the call and let them chill a bit. Sometimes they just need to cool down and call later. Don't worry, they gonna.
Even when you solve a problem they're having, the common reaction is a neutral one (which is totally fine in my book), or even better (tacit or not): "Yeah, you'd better fix it; it wasn't supposed to be broken in the first place!"
It's nice.
non-network people (and some network) dont differentiate between vlans and subnets. if you dont do network and just use it, you usually dont care about l2, you live in the l3 world so it makes more sense just to use ip addresses and subnets. but its probably due to vlan being easier to say and write instead of subnet.
if the routing, firewall policies, config etc. are suboptimal, network users won't feel it in the vast majority of cases; millisecond or two slower due to suboptimal path through the network, or 200 firewall policies instead of 40 optimized ones... but they'll certainly feel it when whole network segment cannot reach some servers because you were too bothered by some routing atrocity you saw and just had to fix it which resulted in rpf fail and aforementioned unreachability. skill issue i suppose.
mids conflate these two.
both the skills and the tribal knowhow of procedures/customs are necessary to excel at your role.
when you can't figure out why some config isnt working, sometimes you have another working config snippet, either on the same device or on thousands of others.. useful trick, but one should strive to be so fuckin good and see the fuckup instantly.
show commands, logs, debug, pcaps. ranked by amount of effort needed to understand the output and ease of setup/execution. the fifth is divine intervention.
if they are a network whiz, they know a bit (at least!) about dev, ops, systems, virtualization, databases, web, security... a devops that has managed a couple of firewalls in their time; active directory guru who made a website or two, a pentester who was a dev etc. those who stick religiously to their narrow niche don't get far. and usually the reason isn't so they can devote more time to their "specialty"; cuz they're mediocre at it as well...
i think it was Nick Russo who said: "you dont have 10 years of experience, you have 10 times the same year of experience". i know some people who have 3 times the "experience", and 3 times less the knowledge. it's okay not to be obsessed with your job/career/craft/calling and continue working/learning even after you've clocked out, but the problem is when you arent at least curious to learn on the job, but look for ways to as little as possible and do the same things in the same way over and over again.
if you came to a maintenance windowsa and think what are all the things i have to change, you've fucked up. you should whiteboard the config, go through all the scenarios, apply the disabled config absolutely certain its gonna work. when the time comes, you just enable and monitor; and then test and monitor. if you prepared as you should, the job is done in the span of minutes. if you haven't.. well, good luck.
speed, efficiency, ease of scripting. use gui if you have to; graphs, filtering, columns, when the cli isn't accessible as it should be... but the goto is the cli. there's a reason some config is available only through the cli, but the other way around, almost never. me, personally, i don't feel i understand the config/solution if i cant configure it through the cli exclusively. and yeah, it should go without mentioning, avoid fuckin wizards; those are fine if you don't plan on being proficient in the topic or at the beginning until you figure out the technology, but the sooner you toss them, the better.
forti docs say about policy routes you can configure either the interface, gateway or both as the destination. turns out it doesnt work if you specify only the interface... oh yeah, rpf looks only at the rib, not the policy routes so set src-check disable.
there is something so satisfying about well written docs. but there is something even more satisfying about figuring out something without docs.
courses, guides, docs, books, papers, experience. wizardrly intuition.
dont know what to think about semi random, semi wizardly knowhow to work around some bug. why tf you cant change non mgmt vlan thats also in a different vrf on mikrotik v6/7 (dont remember anymore), but when you disable everything except the vlan youre connected over, it works... not yet sure if theres method to the madness or is it try changing shit until it works then fucking remember it firmly. prolly both, art and science.