Ending site-wide outages on a CodeIgniter CRM/LMS
A busy CodeIgniter 3 CRM and learning platform kept going down for every user at once. I found that there was not one cause but several, fixed each of them, and validated the result with a 500-virtual-user load test.
Recurring site-wide blackouts
The platform handled attendance, recordings and day-to-day operations for its users, and it went down regularly. When it failed, it failed for everyone at once: pages hung, then the whole site stopped responding until services were restarted.
Restarting brought it back, but the outages kept returning, usually at the busiest times. The goal was not another restart but to find out why it kept happening and make it stop.
Following the evidence
Instead of guessing, I gathered evidence from every layer during normal load and during incidents: PHP-FPM worker status, Apache limits and logs, the MySQL process list and connection counts, firewall behaviour, and slow application endpoints.
That showed the outage was a chain of problems, each making the others worse:
- Undersized PHP-FPM pools: there were too few workers, so requests queued up under normal peak traffic.
- Apache ServerLimit too low for the traffic, capping how many requests could be served at once.
- Firewall rules that interfered with legitimate traffic under load.
- Slow attendance and recording queries that held workers and database connections for far too long.
- Unbounded external API calls with no timeout. When a third-party service was slow, every worker that called it sat waiting, and the pool ran dry.
Fixing each cause, one at a time
I fixed the problems in order of impact, verifying each change before moving on:
- Resized PHP-FPM pools and raised Apache limits to match the server's memory and real traffic.
- Corrected the firewall rules causing problems under load.
- Indexed and rewrote the slow queries behind attendance and recordings, so they returned quickly and released connections.
- Added timeouts to every external API call, so a slow third party could no longer freeze the whole site.
- Moved heavy mobile file uploads into a queue (a DB-backed queue processed by CLI workers under Supervisor), so web requests no longer did the heavy lifting.
Finally, I wrote monitoring scripts that check HTTP health, PHP-FPM worker state, the MySQL process list and connection counts, so any future pressure shows up well before users notice.
Results
Tech stack
- CodeIgniter 3
- PHP-FPM
- Apache
- MySQL
- Supervisor
- Queues
- cPanel/WHM
- Load testing