Crocusoft | Monitoring and Logging: How to Find Production Problems Before Your Users Do
Monitoring dashboard showing latency, error count and server load charts
Technology 5 MIN READ 10/8/2026 6:26:30 AM

Monitoring and Logging: How to Find Production Problems Before Your Users Do

It's six o'clock on a Friday evening. The team is getting ready to head home when a customer calls: "I can't place an order, the payment page won't open." You open the site and it works fine. But the customer is right too. The problem has existed for two hours, and nobody noticed. In the meantime, you'll never know how many people got stuck on the same page, or how many quietly moved on to a competitor's site.

This is every developer's nightmare, and unfortunately it is far more common than it should be. The good news is that there is a fix: properly set up monitoring and logging. In this article we'll cover what they are, how they differ, and how even a small team can put them in place in a simple way.

What Is Logging?

Logging is the diary a system writes about itself while it runs. A user logged in, an order was created, a payment attempt was declined, a connection to the database failed. Every event is recorded as a line with a timestamp.

A log mainly answers one question: what happened? When a problem already exists, the log is the first place you look for the cause. Hunting for a bug without good logs is like searching for a key in a dark room.

What Is Monitoring?

Monitoring means watching the state of your system continuously. How loaded is the server, how many seconds does the site take to respond, how many errors are returned per minute, what share of payments succeed. These are shown as numbers, usually on charts.

Monitoring answers a different question: what is happening right now, and is everything normal? Its main advantage is that it warns you about a problem before it grows.

The Difference Between Logging and Monitoring

CriteriaLoggingMonitoring
Core questionWhat happened?How are things right now?
Data formatLines of text, event recordsNumbers, charts
When you need itWhen investigating a causeTo spot a problem in time
Example"Payment service did not respond for 30 seconds"A sudden rise in the payment failure rate
RoleLike a detective, works out what happenedLike an alarm system, raises the alert

In short, monitoring says "something is wrong," and logging says "here is exactly what is wrong." Having one does not replace the other.

The Third Pillar: Tracing

In a microservices architecture, a single user request passes through several services. When someone places an order, for example, the request goes to the API first, then to the stock service, then to the payment service and the notification service. Searching the logs of each service separately to find where the delay happened takes a lot of time.

Tracing shows that path as a single line: which services the request went through and how much time was lost where. Logs, metrics and traces together are called observability. If you work with a microservices architecture, tracing is no longer a luxury, it's a practical need.

What Should You Monitor?

The most common mistake is trying to monitor everything. You end up with hundreds of charts and nobody looks at any of them. To start, these four indicators are enough:

  • Latency: How long users wait for a response. Look at the slowest requests rather than the average, because unhappy users are usually the ones stuck on those.
  • Error count: How many requests fail per minute. A sudden spike in 500 errors is usually the first sign of a serious problem.
  • Request volume: Is traffic at a normal level, or has it suddenly dropped? A sudden drop in traffic is a problem too, since it can mean users can't reach your site.
  • Resource usage: Server, memory, disk and database connections. When a disk fills up, systems often stop without any warning.

On top of these, always track business indicators as well: how many orders are created per minute, how many payments are completed, how many sign-ups happen. Even if the technical indicators look normal, orders dropping to zero tells you right away where the problem is.

How to Write a Good Log

A few rules make the difference between a useful log and a useless one.

  • Use levels correctly: Info for normal events, Warning for suspicious situations, and Error only for operations that truly failed. If everything is logged as Error, you can't find the real error.
  • Add context: "Payment failed" isn't enough. Say which order, which user, what reason and what time.
  • Use a structured format: Logs written in JSON are very convenient for searching and filtering.
  • Don't write sensitive data: Passwords, card numbers and tokens should never end up in a log. This is a serious security risk, and one of the issues we touched on in our API security article.
  • Collect everything in one place: If your system runs on several servers, logs should be gathered centrally. Logging into each server separately during a real incident is wasted time.

Alerts: Let the Problem Tell You Itself

Nobody can stare at charts all day. That is why the most important part of monitoring is alerts: an automatic warning sent when a certain condition is broken. For example, "if the error rate exceeds seven percent over the last five minutes, message the team."

The key rule when setting up alerts is this: every alert should require a person to act. If nothing needs to be done when an alert arrives, that alert is redundant. Too many unnecessary alerts train the team to ignore warnings, and as a result a truly important notification gets missed.

It helps to split alerts into two groups. Urgent ones (the site is down, payments aren't going through) should reach people immediately by call or message. Non-urgent ones (the disk is at 70 percent) should land in a list that gets reviewed during the day.

Monitoring After a Deployment

A big share of problems appear right after a new version is released. That's why monitoring should be thought of together with your CI/CD pipeline. Watching the error count and latency in the first minutes after a new version reaches production, and rolling back to the previous version immediately if there's a problem, is one of the best habits a team can have.

Some teams take this even further: a new version is first shown to a small share of traffic, and if the metrics stay normal it is gradually opened to everyone. This approach depends on monitoring being set up properly, and it cuts risk significantly.

Monitoring in Container and Kubernetes Environments

Containers are temporary by nature. When a Pod restarts, the files inside it, including log files, can be lost. That's why in a Kubernetes environment, logs should not be kept inside the container but sent to a central system. You also need to watch resource usage at the Pod, Node and Cluster level, because sometimes the problem isn't in the application itself but in the application not getting enough resources.

Which Tools Can You Use?

There are plenty of options on the market, and which one fits depends on the size of your project.

ToolWhat it's known for
Prometheus and GrafanaOpen source metric collection and dashboards, very common with Kubernetes
ELK (Elasticsearch, Logstash, Kibana)Collecting, searching and visualizing logs
LokiA relatively lightweight log system that works well with Grafana
OpenTelemetryCollecting metrics, logs and traces in a standard format
SentryAutomatically catches application errors and notifies the team
Uptime monitorsCheck from the outside whether your site is reachable

For a small project, the simplest start is an uptime monitor that checks the site from the outside, a tool like Sentry that catches application errors, and one dashboard showing the core indicators. Even these three remove most of the "learning about the problem from a customer" situations.

The Most Common Mistakes

  • Leaving monitoring for later: The "let it work first, we'll set it up afterward" approach usually lasts until the first serious incident.
  • Only watching the server: The server can look healthy while the payment service returns errors. Business indicators need watching too.
  • Setting up too many alerts: When everything sends a warning, none of it is taken seriously.
  • No log retention policy: Logs grow fast. If you don't decide how long to keep them, the disk fills up and becomes a problem of its own.
  • Not assigning responsibility: It should be clear in advance who looks at an alert when it arrives.
  • Only monitoring Production: Staging needs monitoring too, so problems get caught before they reach users.

How to Get Started

You don't need to build everything in a day. This order works well on real projects:

  1. Set up a simple uptime monitor that checks your site from the outside.
  2. Add a tool that automatically collects application errors.
  3. Gather logs in one central place and move to a structured format.
  4. Create one dashboard for the four core indicators (latency, errors, volume, resources).
  5. Set up two or three of the most important alerts and decide who is responsible.
  6. Add business indicators (orders, payments) to the dashboard.
  7. Add tracing as the system grows.

If your project is already complex, or your team has no experience in this area, building the monitoring and log architecture correctly from day one costs far less than fixing it later. Work like this is usually considered at the architecture stage of custom software projects.

Frequently Asked Questions

Does a small project need monitoring?
Yes. Even the simplest uptime check tells you before your customers do when the site goes down. It costs almost nothing.

Are logging and monitoring the same thing?
No. Monitoring shows the current state in numbers and warns you about problems. Logging records in detail what happened. They complement each other.

Is it right to log every event?
No. Too much logging raises costs and buries the important information. It's better to log the key events and every error, with context.

What should you do when an alert comes in at night?
There should be a simple, pre-written runbook: what to check, who to call, and how to roll back if needed. At three in the morning, nobody finds it easy to remember all of that.

Does monitoring affect performance?
When set up properly, the impact is very small. The main risk is writing overly detailed logs.

Conclusion

You can't promise that problems will never happen. But whether you hear about them first is up to you. Monitoring warns you about a problem, logging helps you find the cause, and tracing shows the way through complex systems. You don't need a big budget to set them up. You just need the right order of steps and the decision to start.

If you want to set up monitoring and logging for your project, or strengthen what you already have, get in touch with the Crocusoft team. Together we can work out which step makes sense to start with.