EN
Webmail

Server Log Analysis: Diagnosing Website Problems From Access Logs

Server Log Analysis: Diagnosing Website Problems From Access Logs

Server log analysis is the most direct way to understand what is really happening on a website. Analytics tools show what visitors do in their browsers, and only when their browsers run the tracking script. The web server’s logs record every single request: real visitors, search engine crawlers, monitoring services, bots probing for weaknesses and requests that fail before any page loads. When a site is slow, throws errors, gets attacked or mysteriously drops out of search, the logs usually contain the answer.

This guide explains what access and error logs contain, how to read them, and how to use them to diagnose the problems we most often see in server administration work: broken URLs, server errors, abusive bots, slow requests and crawling issues. It is written for developers, website owners and marketers who want to use logs without becoming full-time system administrators.

What Server Logs Contain

Most websites run on Apache or Nginx, often behind a CDN or load balancer. Both web servers write two main kinds of logs.

Access Logs

The access log records one line per request. In the widely used combined format, described in the Apache logging documentation, each line contains:

  • The client IP address.
  • The date and time of the request.
  • The request method and path, for example GET /pricing/ HTTP/1.1.
  • The HTTP status code returned, such as 200, 301, 404 or 500.
  • The size of the response in bytes.
  • The referrer, meaning the page that linked to this request, if the browser sent it.
  • The user agent string, which identifies the browser or bot.

A typical line looks like this:

203.0.113.7 - - [06/Oct/2026:09:14:22 +0000] "GET /blog/ HTTP/2.0" 200 18342 "https://www.google.com/" "Mozilla/5.0 ..."

Many administrators extend the format with the request processing time, which is invaluable for performance work. Nginx’s log module can record $request_time and upstream response times, and Apache can log %D for the time taken in microseconds.

Error Logs

The error log records problems the server or application encountered: PHP fatal errors and warnings, failed connections to the database or an upstream service, permission problems, configuration errors and timeouts. Access logs tell you that a request returned 500; the error log usually tells you why.

Application and CDN Logs

Applications such as WordPress or e-commerce platforms can write their own logs, and CDNs keep logs of requests they answered from cache without contacting your server. When a CDN is in use, remember that your server’s access log only sees requests that passed through to the origin, and the client IP may be the CDN’s address unless the real IP is restored from a header.

Getting Started Safely

Logs contain personal data such as IP addresses, so treat them accordingly: limit access, keep them only as long as needed and avoid copying them to random laptops. On a typical Linux server, logs live under /var/log/apache2/, /var/log/httpd/ or /var/log/nginx/, and hosting control panels often keep separate logs per website. Logs are usually rotated daily or weekly and compressed, so older periods sit in files ending in .gz.

For quick investigations, standard command-line tools are enough. For ongoing analysis, a log analyser or a central logging system that indexes logs from several servers saves a great deal of time. Either way, the questions you ask are the same.

Five Problems Logs Help You Diagnose

ProblemWhat to look for in logsTypical fix
Broken URLsFrequent 404 responses, grouped by path and referrerRedirect moved pages, fix internal links
Server errors5xx responses in the access log, matching entries in the error logFix the failing code, resource limit or dependency
Abusive botsMany requests from few IPs or user agents, login and admin probesRate limiting, firewall rules, blocking
Slow pagesHigh request times by URL and time of dayCaching, query optimisation, more resources
Crawling issuesWhich URLs verified search bots fetch, and their status codesFix redirects, errors and crawl traps

1. Finding Broken URLs

Group 404 responses by requested path and count them. A handful of paths usually accounts for most of the errors. Then look at the referrer for each one:

  • Referrer is your own site: an internal link is broken. Fix the link at the source.
  • Referrer is another website: someone links to a page that moved. Add a redirect to the correct page so you keep the visitor and the link’s value.
  • No referrer, requested by bots: often old URLs that search engines remember, or random probes. Redirect the ones that had real content; ignore the probes.

Redirect chains also show up here: a request that returns 301, followed by another 301 and finally 200. Each hop wastes time and crawl effort, so point links and redirects directly at the final address.

2. Diagnosing Server Errors

Filter the access log for status codes from 500 to 599 and note the times and paths. Then open the error log for the same minutes. Common patterns include:

  • 500 Internal Server Error with a PHP fatal error, often after a plugin or code update.
  • 502 Bad Gateway when the web server cannot get a response from PHP-FPM or an application server, often because worker processes are exhausted or crashed.
  • 503 Service Unavailable during maintenance or when a protection layer is rejecting requests.
  • 504 Gateway Timeout when the application takes too long, frequently because of slow database queries or a slow external API.

Errors clustered at specific times point to scheduled jobs, backups or traffic peaks. Errors on a single URL point to a code problem on that page.

3. Spotting Abusive Bots

Count requests by IP address and by user agent for a recent period. Legitimate traffic is spread across many addresses. Abuse often looks like one address making thousands of requests, repeated hits on login pages such as /wp-login.php or /xmlrpc.php, or probes for files that do not exist on your platform, such as configuration backups or admin panels for other software.

Responses range from rate limiting and blocking specific addresses to a web application firewall that filters known attack patterns. Be careful not to block legitimate crawlers or monitoring services; check what an address is before blocking it. If you suspect a site has already been compromised, our guide to website malware and blacklisting covers detection and recovery.

4. Finding Slow Requests

With request times in the log, you can sort by duration and group by path. This reveals which pages are slow for real users, not just in a single test. Look for:

  • Pages that are consistently slow, which need code, query or caching improvements.
  • Pages that are slow only at certain times, which suggests resource contention or scheduled jobs.
  • Uncached requests that should be cached, such as anonymous visits to public pages.
  • Heavy endpoints called very often, such as search or AJAX handlers.

Combining log data with front-end metrics gives a complete performance picture: logs show server time, while browser metrics show what visitors experience after the response arrives.

5. Seeing What Search Engines Crawl

Search Console shows Google’s view of your site, but logs show exactly which URLs Googlebot requested and what your server returned. Filter for search engine user agents, then verify that the requests really come from the search engine, because many bots fake the Googlebot user agent. Google explains how to verify Googlebot with a reverse DNS lookup or its published IP ranges.

With verified crawler requests, you can answer useful questions: Are important pages crawled regularly? Is crawl effort wasted on parameter URLs, filters or internal search? Do crawlers hit redirects and errors? If pages are missing from search, these findings complement the checks in our article on why Google is not indexing your pages.

A First Look With Simple Commands

You do not need special software to start. On a Linux server with a combined-format access log, a few standard commands answer the most common questions. Replace the file name with your own log.

Count responses by status code:

awk '{print $9}' access.log | sort | uniq -c | sort -rn

List the twenty most requested paths that returned 404:

awk '$9 == 404 {print $7}' access.log | sort | uniq -c | sort -rn | head -20

Find the IP addresses with the most requests:

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -20

Show requests whose user agent claims to be Googlebot, so you can verify them:

grep -i "googlebot" access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head

Field positions depend on the exact log format, so check one line first and adjust the numbers if your format differs. For compressed older logs, use zcat or zgrep instead of cat and grep.

These commands are enough to answer “what is going wrong right now?” in a few minutes. When you need trends over weeks, comparisons between servers or dashboards for a team, move to a log analyser or a central logging platform.

Correlating Logs With Other Data

Logs become even more useful when you line them up with other information. A spike in 5xx errors that starts exactly when a deployment finished points to the release. Slow responses that coincide with a backup job point to resource contention. A drop in crawler activity a few days after a robots.txt change points to the new rules. Keep a simple change log of deployments, configuration changes, plugin updates and campaigns with dates and times. When something unusual appears in the server logs, that change log is often the fastest route to the cause.

Making Log Analysis a Routine

Logs are most valuable when someone looks at them before a crisis. A light routine for a business website:

  1. Daily, automated: alerts for spikes in 5xx errors and for unusual request volumes.
  2. Weekly: a short report of the top 404 paths, top error messages and slowest URLs.
  3. Monthly: a review of crawler activity and of the heaviest IP addresses and user agents.
  4. After every release: a check of error logs for new warnings and errors.

Make sure logs are rotated and retained for a defined period, that disk space is monitored so logs never fill the server, and that the time zone is consistent across servers so events can be correlated. Our server administration service keeps client servers updated, protected and backed up, with remote support when something needs attention.

Frequently Asked Questions

What is the difference between access logs and error logs?

Access logs record every request and its result. Error logs record problems the server or application encountered, which explains why some requests failed.

Why do my analytics numbers differ from my server logs?

Analytics only counts browsers that run the tracking script and accept it. Logs count every request, including bots, blocked scripts and assets such as images.

How long should I keep server logs?

Long enough to investigate incidents and trends, often a few weeks to a few months, and no longer than your privacy policy and legal obligations allow, since logs contain IP addresses.

Can I analyse logs on shared hosting?

Often yes. Many hosting control panels provide raw access logs for download, although error logs and custom formats may be limited.

How do I know if a request really came from Googlebot?

Run a reverse DNS lookup on the IP address, check that it belongs to Google’s domains and confirm with a forward lookup, or compare against Google’s published IP ranges.

Do I need special software for log analysis?

Not for occasional checks. Command-line tools work well. For regular analysis across several sites or servers, a central logging system or log analyser saves time.

The Bottom Line

Server log analysis shows the complete picture that analytics misses: every request, every error and every crawler visit. Use access logs to find broken URLs, abusive bots, slow pages and crawling problems, and error logs to explain the server errors you find. Add request times to your log format, verify search engine crawlers before trusting them, protect logs as personal data and turn the checks into a regular routine with automated alerts. Problems found in logs are usually cheaper to fix than problems found by customers. If you want help setting up logging or investigating an issue, contact our team.