Outage tracing
Alarms and complaints become one incident, with the exact list of affected customers and the most likely location of the fault. After the repair, Eldaven re-runs the checks and closes the incident only when every customer is back.
In development since April 2026, inside a working ISP
Eldaven is outage operations software for local ISPs in Indonesia. Machine learning reads every alarm from your MikroTik and fiber network, groups them into incidents and writes the first report. A language model is called in only when an incident needs real reasoning.
Simulated network. In a real deployment, Eldaven builds this map from your routers, OLT and customer records.
I run a small internet provider in Indonesia. For years, every outage started the same way: customers messaging us on WhatsApp before our team knew anything was wrong.
In April 2026 I started building tools to fix that on our own network. Most local ISPs run exactly like we do, so we're turning those tools into a product.
Small, fast models do the volume work on your own server. When an incident doesn't fit a known pattern, Eldaven can hand it to the LLM of your choice for a deeper look.
Of 0 alarms, 0 became incidents and 0 needed a language model. That's 0% of alarms. Simulated; real ratios depend on your network.
Cost and speed. A small ISP sees thousands of alarms a day. Sending each one to a language model would cost more per month than many customers pay, and every answer takes seconds. ML models run locally in milliseconds with no usage fee, so they carry the load.
Some incidents don't match any pattern: a fault after a config change, or symptoms spread across sites. There, a language model can read logs and history, compare explanations, and write them up in plain Bahasa Indonesia. It's a tool for the hard 1%, not the routine 99%.
| ML engineAlways on | Language modelOptional, on demand | |
|---|---|---|
| Does | Detects unusual traffic, signal and session patterns. Groups alarms by cable. Runs standard read-only checks. Writes the incident note. | Investigates unusual faults across logs and config history. Explains its reasoning. Drafts updates for customers. |
| Runs | On every alarm, around the clock. | When ML marks an incident as unclear, or a technician asks. |
| Lives | On your server, next to your network. | Your provider and your API key: Claude, Gemini, GPT, or an open model on your own hardware. |
| Costs | No usage fee. | Billed by your provider, capped by a monthly limit you set. |
Once Eldaven knows which customer sits behind which cable, that knowledge is useful far beyond the NOC: to CS, to field technicians, and to whoever plans the next upgrade.
Alarms and complaints become one incident, with the exact list of affected customers and the most likely location of the fault. After the repair, Eldaven re-runs the checks and closes the incident only when every customer is back.
Every customer's path from POP to ODC, ODP and port. Click any device to see whose service depends on it.
First versionYour best technician's know-how as checks anyone can run: customer offline, PPPoE drops, slow at peak hours, DNS trouble, problems after a config change.
First versionLook up a customer and see at once whether they're part of an outage, with a reply that doesn't promise a repair time nobody knows yet.
NextThe location, the likely ODP, the readings to take and the customers to compare against. Results are logged from the technician's phone.
NextRouter configuration snapshots compared over time, so you can see what changed right before something broke.
NextThe next shift sees what was checked and what's left. Confirmed causes are saved, so the same fault is solved faster the second time.
NextForecasts when peak-hour traffic will reach your uplink limit, so you can upgrade before customers feel it.
LaterBefore scheduled work, see which customers will be affected and which checks to run afterwards.
LaterNo new hardware and no change to how your network is configured. Eldaven connects with read-only access to what you have.
| Area | Integration | How it connects | Status |
|---|---|---|---|
| Routers | MikroTik RouterOS 6 and 7 | API, REST and SNMP, through a read-only user | First version |
| Other SNMP devices | SNMP v2c and v3 | Later | |
| Fiber | SFP optical readings | Rx and Tx power from MikroTik SFP ports | First version |
| OLT and ONT signal | One OLT vendor first, chosen with pilot ISPs | Next | |
| Customers | PPPoE sessions and DHCP leases | Read from your MikroTik routers | First version |
| Customer list | CSV import from your billing system | First version | |
| RADIUS and billing systems | Direct integration | Later | |
| Alerts | Email and Telegram | Incident notifications to your team | First version |
| Updates for your CS team | Next | ||
| Language models | None, ML only | Everything stays on your server | First version |
| Claude, Gemini, GPT | Your own API key, with a monthly spending limit | First version | |
| Open models | Running on your own hardware | Later |
Customer data and router access are the most sensitive things an ISP has. These are the rules Eldaven is being built around.
We move to the next stage only when the current one works on a real network.
First tools built for our own ISP.
Collector, alarm grouping and the first ML models, built on our own network.
Two or three outside ISPs, five checks for common faults, optional LLM investigations.
Billing and RADIUS links, OLT signal data, capacity forecasts, multi-site.
We're looking for two or three local ISPs to shape the first version with us. Tell us which outages cost your team the most time.
hezron@eldaven.online