Implementing MikroTik Watchdog Auto Recovery: Solving Fiber Optic Network Issues After Power Outages
Implementasi Watchdog MikroTik: Solusi Pemulihan Jaringan FO Setelah Gangguan Listrik
Pernah ngalamin kejadian listrik mati, terus pas nyala lagi jaringan internet malah nggak balik normal?
Jadi ceritanya gini.. di ruang server ada kejadian MCB turun. UPS udah mati total karena listrik nggak balik-balik. Terus pas listrik nyala, server sama MikroTik otomatis nyala. Server sih aman karena auto power-on. Tapi masalahnya.. jaringan Fiber Optic (FO) nggak langsung nyambung. Padahal kabel udah nyala, port udah up, tapi trafik nggak jalan.
Di kejadian sebelumnya, satu-satunya cara biar FO normal lagi adalah dengan reboot MikroTik secara manual. Ya gitu deh, harus ada orang yang dateng ke ruang server atau akses remote buat reboot. Ribet banget kalau kejadiannya malam atau pas lagi nggak ada orang.
Nah, dari situ muncul ide: harus ada mekanisme otomatis yang bisa mendeteksi kalau FO bermasalah dan langsung reboot MikroTik sendiri tanpa nunggu manual.
Perangkat dan Kondisi Awal
MikroTik yang dipakai adalah CCR1016-12G dengan RouterOS versi 6.48.6 long-term. Perangkat ini lumayan tua tapi masih andal banget buat jadi router utama.
Server di LAN sebenarnya bisa di-ping, tapi diputuskan server cuma jadi indikator tambahan atau log aja. Kenapa? Karena kalau server mati tapi FO masih jalan, MikroTik nggak perlu reboot. Logikanya gitu.
Penentu utama reboot adalah konektivitas FO melalui beberapa host dari VLAN 60, VLAN 70, dan VLAN 90. Tiga VLAN ini mewakili segmen jaringan FO yang berbeda. Kalau salah satu aja bisa di-ping, berarti FO sehat.
Logika Dasar Auto Recovery
Gimana mekanismenya? Sederhana sebenernya:
- Minimal satu host dari VLAN 60, 70, atau 90 harus reachable. Kalau iya, FO dianggap sehat, nggak usah reboot.
- Kalau semua host gagal di-ping, tunggu 60 detik, terus cek lagi.
- Di pengecekan kedua, kalau salah satu VLAN balik aktif, reboot dibatalkan.
- Kalau tetep semua gagal, reboot MikroTik.
Tapi ada yang lebih penting lagi.
Anti Reboot Loop
Nah, ini yang sering dilupain. Kalau setelah reboot terus FO tetep nggak normal, MikroTik nggak boleh reboot terus-terusan. Jadi dipasang recovery lock berupa file penanda.
Gini cara kerjanya:
- Sebelum reboot, script bikin file lock.
- Kalau lock udah ada, reboot kedua diblokir.
- Kalau FO kembali normal, lock dihapus.
Jadi maksimal cuma satu kali reboot otomatis. Kalau setelah itu masih mati, ya udah, berarti ada masalah lain yang perlu dicek manual.
Pengujian: Sebelum Diproduksi
Nggak mungkin langsung dipasang ke produksi kan. Jadi dibuat script buat testing dulu.
TEST-FO-HEALTH-CHECK-V3
Script ini buat ngetes health check doang. Ada delay awal 5 menit (biar nggak langsung jalan pas booting), terus cek host-host VLAN. Tapi nggak di-reboot, cuma log aja hasilnya. Hasilnya? Logika minimal satu VLAN FO aktif bekerja sesuai harapan.
TEST-DECISION-MATRIX
Ini lebih kompleks. Ngetes kombinasi kondisi server dan tiga VLAN FO. Misalnya server hidup semua VLAN mati, atau server mati cuma satu VLAN hidup, dan seterusnya. Semua keputusan sesuai matrix yang udah ditentukan. Berhasil semua.
Script Produksi: FO-AUTO-RECOVERY
Setelah testing oke, dibuat script produksi dengan alur lengkap:
- RouterOS boot → script otomatis dijalankan
- Tunggu 5 menit (biar semua service stabil dulu)
- Health Check #1 → cek semua VLAN FO
- Kalau minimal satu VLAN OK → nggak reboot, selesai
- Kalau semua FO gagal → tunggu 60 detik
- Health Check #2 → cek lagi
- Kalau ada yang pulih → nggak reboot
- Kalau tetap gagal → cek recovery lock
- Lock belum ada → bikin lock, reboot
- Lock udah ada → blokir reboot kedua, log aja
Scriptnya ditulis di System → Scripts dan di-save dengan nama FO-AUTO-RECOVERY.
Biar Otomatis: Scheduler Startup
Script udah jadi, tapi harus dijalankan otomatis setiap MikroTik boot. Makanya dibuat scheduler dengan nama FO-AUTO-RECOVERY-STARTUP.
Konfigurasinya:
- start-time=startup (jalan pas booting)
- interval=0s (cuma sekali)
- event=FO-AUTO-RECOVERY (jalankan script)
Udah dibuat, udah aktif, dan diverifikasi. Kalau MikroTik mati terus nyala lagi, scheduler ini langsung menjalankan script FO-AUTO-RECOVERY.
Nggak Ngaruh ke Konfigurasi Lain
Oh iya, script queue yang udah ada sama scheduler HTTP check/failover ISP dibiarkan aja. Soalnya kebutuhan recovery FO ini terpisah dan nggak bersinggungan sama konfigurasi existing. Nggak ada perubahan di bagian itu.
Dokumentasi dan SOP
Setelah semua jadi, disusun juga SOP pengecekan buat tim operasional. Ini penting banget, apalagi kalau kejadian berulang dan ada yang perlu dicek manual.
Yang dicek:
- Uptime CCR — untuk mastiin router udah restart atau belum
- Scheduler — masih aktif nggak
- Log FO-AUTO-RECOVERY — liat riwayat health check
- Recovery lock — ada atau nggak
- Status server — sebagai indikator tambahan
- Konektivitas VLAN 60/70/90 — dari sisi client
Lokasi konfigurasi di GUI juga didokumentasikan:
- System → Scripts
- System → Scheduler
- Files (tempat recovery lock)
- Log
Dokumentasi ini kemudian dibuat dalam bentuk PDF sebagai catatan operasional.
Usulan Buat Atasan
Setelah semuanya jalan dan terbukti, konsep ini dirumuskan sebagai usulan ke atasan. Intinya: mekanisme watchdog/auto recovery ini solusi untuk mengurangi kebutuhan reboot manual kalau kejadian MCB turun dan FO nggak normal terulang lagi.
Selain itu, ini juga jadi bukti kalau tim IT punya inisiatif buat nge-prevent masalah sebelum terjadi, bukan cuma nunggu masalah terus nge-fix manual.
Refleksi
Gw pribadi cukup puas sama hasilnya. Prosesnya mulai dari identifikasi masalah, bikin logika, testing, sampe implementasi produksi itu sebuah rangkaian yang utuh. Nggak instan, tapi worth it banget.
Satu hal yang gw pelajari: jaringan itu harus bisa sembuh sendiri, terutama buat infrastruktur kritikal. Ketergantungan sama intervensi manual itu risiko. Dan risiko itu harus diminimalisir.
Nah, kalau ada yang lagi nyari solusi buat masalah serupa, semoga ini bisa jadi referensi. Intinya:
- Cari tahu apa indikator sehat jaringanmu
- Bikin logika yang cukup sederhana tapi reliable
- Jangan lupa anti reboot loop
- Testing dulu sebelum produksi
- Dokumentasi biar gampang nge-track
Ya mungkin segitu dulu. Semoga bermanfaat!
Artikel ini adalah catatan teknis pribadi dari pengalaman implementasi Watchdog MikroTik untuk pemulihan otomatis jaringan Fiber Optic.
Implementing MikroTik Watchdog: Solving Fiber Optic Network Recovery After Power Outages
Ever had a power outage, and when the electricity came back, your internet just wouldn't return to normal?
So here's what happened.. there was a power trip in the server room. The UPS had completely drained because power didn't come back quickly. When power was restored, the servers and MikroTik turned on automatically. The servers were fine because they have auto power-on. But the problem? The Fiber Optic (FO) network didn't come back online. The cables were lit, the ports were up, but traffic just wouldn't flow.
In previous incidents, the only way to restore FO connectivity was to manually reboot the MikroTik. That meant someone had to physically go to the server room or access it remotely. Really inconvenient if it happens at night or when no one's around.
That's where the idea came from: we need an automatic mechanism that can detect FO issues and reboot the MikroTik itself without waiting for manual intervention.
Hardware and Initial Setup
The MikroTik in use is a CCR1016-12G running RouterOS version 6.48.6 long-term. It's an older device but still incredibly reliable as a core router.
The LAN server could technically be pinged, but we decided to use the server only as a supplementary indicator or log. Why? Because if the server goes down but the FO is still running, the MikroTik doesn't need a reboot. That's the logic.
The primary trigger for reboot is FO connectivity through hosts on VLAN 60, VLAN 70, and VLAN 90. These three VLANs represent different FO network segments. If just one of them is reachable, the FO is considered healthy.
The Auto Recovery Logic
How does the mechanism work? It's actually quite simple:
- At least one host from VLAN 60, 70, or 90 must be reachable. If yes, FO is healthy — no reboot needed.
- If all hosts fail the ping test, wait 60 seconds, then check again.
- During the second check, if any VLAN becomes active again, the reboot is canceled.
- If all still fail, reboot the MikroTik.
But there's something even more important.
Anti Reboot Loop
This is often overlooked. If the FO still doesn't recover after the reboot, the MikroTik shouldn't keep rebooting endlessly. So we implemented a recovery lock — a marker file.
Here's how it works:
- Before rebooting, the script creates a lock file.
- If a lock already exists, the second reboot is blocked.
- If the FO returns to normal, the lock is removed.
So there's a maximum of one automatic reboot. If the FO is still down after that, there's probably a different issue that requires manual investigation.
Testing Before Production
We couldn't just deploy this directly to production. So we created scripts for testing first.
TEST-FO-HEALTH-CHECK-V3
This script was designed to test only the health check logic. It has an initial 5-minute delay (to avoid running immediately after boot), then checks all VLAN hosts. But it doesn't actually reboot — it just logs the results. The result? The "at least one VLAN active" logic worked as expected.
TEST-DECISION-MATRIX
This was more comprehensive. It tested combinations of server conditions and the three FO VLANs. For example: server up but all VLANs down, server down but only one VLAN alive, and so on. All decisions matched the predetermined matrix. Everything passed.
Production Script: FO-AUTO-RECOVERY
After successful testing, we created the production script with this full workflow:
- RouterOS boots → script runs automatically
- Wait 5 minutes (let all services stabilize)
- Health Check #1 → check all FO VLANs
- If at least one VLAN is OK → no reboot, exit
- If all FO checks fail → wait 60 seconds
- Health Check #2 → check again
- If any recover → no reboot
- If still all fail → check recovery lock
- Lock doesn't exist → create lock, reboot
- Lock already exists → block second reboot, just log
The script is stored in System → Scripts and saved as FO-AUTO-RECOVERY.
Making It Automatic: Startup Scheduler
The script was ready, but it needed to run automatically every time the MikroTik boots. So we created a scheduler named FO-AUTO-RECOVERY-STARTUP.
Configuration:
- start-time=startup (runs during boot)
- interval=0s (runs only once)
- event=FO-AUTO-RECOVERY (executes the script)
It's created, active, and verified. If the MikroTik goes down and comes back up, this scheduler triggers the FO-AUTO-RECOVERY script immediately.
No Impact on Other Configurations
Oh by the way, the existing queue scripts and ISP failover HTTP check schedulers were left untouched. This FO recovery requirement is separate and doesn't interfere with existing configurations. No changes were made there.
Documentation and SOP
After everything was implemented, we also prepared a Standard Operating Procedure (SOP) checklist for the operations team. This is really important, especially if the issue recurs and requires manual verification.
Items to check:
- CCR uptime — to see if the router has restarted or not
- Scheduler — still active?
- FO-AUTO-RECOVERY logs — check the health check history
- Recovery lock — present or not?
- Server status — as a supplementary indicator
- VLAN 60/70/90 connectivity — from the client side
Configuration locations in the GUI were also documented:
- System → Scripts
- System → Scheduler
- Files (where the recovery lock is stored)
- Log
This documentation was then compiled into a PDF as operational notes.
Proposal for Management
After everything was running successfully, this concept was formalized as a proposal to management. The core message: this watchdog/auto recovery mechanism is a solution to reduce manual reboot needs if power trips and FO issues occur again.
Additionally, it serves as proof that the IT team takes proactive measures to prevent problems, rather than just waiting for issues and fixing them manually.
Final Thoughts
Personally, I'm quite satisfied with the outcome. The process — from problem identification, logic design, testing, to production implementation — was a complete, cohesive workflow. It wasn't instant, but it was definitely worth it.
One key lesson learned: networks should be self-healing, especially critical infrastructure. Relying on manual intervention is a risk. And that risk must be minimized.
For anyone looking for a solution to a similar problem, hopefully this can serve as a reference. In summary:
- Identify what indicates a healthy network
- Design logic that's simple yet reliable
- Don't forget the anti-reboot-loop mechanism
- Test thoroughly before going live
- Document everything for easier tracking
I guess that's all for now. Hope this helps!
This article is a personal technical note based on real-world implementation of a MikroTik Watchdog for automatic Fiber Optic network recovery.
Terima kasih sudah mampir! Jika kamu menikmati konten ini dan ingin menunjukkan dukunganmu, bagaimana kalau mentraktirku secangkir kopi? 😊 Ini adalah gestur kecil yang sangat membantu untuk menjaga semangatku agar terus membuat konten-konten keren. Tidak ada paksaan, tapi secangkir kopi darimu pasti akan membuat hariku jadi sedikit lebih cerah. ☕️
Thank you for stopping by! If you enjoy the content and would like to show your support, how about treating me to a cup of coffee? �� It’s a small gesture that helps keep me motivated to continue creating awesome content. No pressure, but your coffee would definitely make my day a little brighter. ☕️ Buy Me Coffee

Post a Comment for "Implementing MikroTik Watchdog Auto Recovery: Solving Fiber Optic Network Issues After Power Outages"
Post a Comment
You are welcome to share your ideas with us in comments!