OfficeCore builds OfficeTrack, a platform for managing mobile workforces and vehicle fleets, used by more than a thousand business customers. Between 51 and 200 nodes carry it across every environment, and the team looking after them is between two and five people.
Visibility that had thinned, and a bill that had not
Two problems arrived together, which is often how monitoring decisions get made.
“Visibility was not always the best in the past, and price got crazy high over time.”
Dani Avni
CTO, OfficeCore
Neither on its own forces a change. A gradually rising bill is tolerable while the tool works, and imperfect visibility is tolerable while the bill is reasonable. It is the combination that starts an evaluation. OfficeCore chose Netdata for per-second real-time granularity, predictable cost, ML anomaly detection, and the breadth of integrations.
An honest note on the migration
Replacing a monitoring stack that a team has shaped around itself for years is not a drop-in exercise, and OfficeCore’s account of it is refreshingly unvarnished.
“It took some ironing of issues to align with what we had, but now that it all runs we are almost at parity with what we had. We are now able to analyze things using the AI that were not possible before to find the root cause.”
Dani Avni
CTO, OfficeCore
Both halves of that are worth reading. Reaching parity took work — every accumulated check, threshold and habit has to be reproduced or deliberately dropped. What sits on top of parity is the part that was not available at any price before.
Where Netdata AI changed the work
“Finding the root cause has dropped in some cases to minutes.”
Dani Avni
CTO, OfficeCore
The incident behind that is a good illustration of a specific failure mode: when you can only see the part of a problem that reaches customers, you fix the part that reaches customers.
One visible symptom, three crashed services, one shared cause.
“Just last week a service crashed on one of our machines. When asking Netdata AI to analyze it, the AI saw that more than one service actually crashed — we only saw the one that affected customers. It later turned out that there was a memory leak in a 3rd party component that caused the machine to run out of memory, which caused the services to crash. Without the AI fast analysis I would probably have figured it out, but at a slower rate, so I might have lost logs or other information that was still available because of the fast AI analysis.”
Dani Avni
CTO, OfficeCore
That last sentence describes something rarely stated plainly: evidence has a shelf life. Logs rotate, buffers overwrite, processes restart and take their state with them. An investigation that takes hours is not simply slower than one that takes minutes — it runs against a clock, and the data it most needs may not survive to the end of it. Arriving early enough to still have the evidence is a different outcome from arriving late and reconstructing.
The other detail is the count. The team saw one crash because one crash reached customers. The AI saw three, and three crashes on one machine point somewhere a single crash does not: not at the service, but at the host. That reframing is what led to the memory exhaustion, and from there to the third-party component leaking.
What it adds up to
- Root cause in minutes, in cases that used to run long enough to put the evidence at risk.
- The whole failure, not the visible part of it: three crashed services surfaced where the team had seen one.
- Evidence reached while it still exists: fast analysis arrives before logs rotate and state is lost.
- A third-party leak identified: the cause sat in a component the team did not write, which is the hardest kind to find by inspection.
- Predictable cost, replacing a bill that had climbed steadily.
Discover how Netdata can elevate your operations. Learn more.








