Manajemen Bandwidth di Jaringan Rumah Sakit: Studi Kasus 13 VLAN dan Dual ISP
Manajemen Bandwidth di Jaringan Rumah Sakit: Studi Kasus 13 VLAN dan Dual ISP
Setelah memahami prinsip granularitas menggunakan mangle dan queue tree, mari kita lihat bagaimana prinsip tersebut diterapkan—atau justru disederhanakan—dalam sebuah lingkungan dengan kebutuhan nyata namun arsitektur routing yang relatif straightforward.
Jadi, begini ceritanya. Ada sebuah rumah sakit swasta menengah di pinggiran kota. Mereka punya 13 VLAN. Bukan jumlah yang main-main. VLAN-VLAN itu mewakili dunianya masing-masing: ada VLAN untuk Manajemen, IT, Keuangan, HRD, Rekam Medis, Radiologi, Laboratorium, Farmasi, Rawat Jalan, Rawat Inap, Gizi, Engineering, dan yang terakhir—tentu saja—Guest. Setiap departemen punya mikrokosmiknya sendiri, dengan server khusus, printer khusus, dan tentu saja, keluhan internet yang khusus pula.
Koneksi internetnya dual ISP. Satu fiber optik 100Mbps dari provider besar sebagai *primary*. Satunya lagi, radio 50Mbps dari provider lokal sebagai *backup*. Konfigurasinya active-standby murni. Tidak ada load-balancing rumit. Jika fiber mati, semua lalu lintas dialihkan ke radio. Sederhana, bisa diprediksi. Ini poin penting. Karena routing-nya sederhana, logika manajemen bandwidth-nya pun bisa ikut sederhana. Di sinilah pertama kali kita belajar: jangan paksa solusi kompleks jika arsitektur dasarnya sudah sederhana.
Kebutuhan utamanya jelas: pastikan lalu lintas kritis rumah sakit—seperti akses ke aplikasi lab, radiologi digital (PACS), dan rekam medis elektronik—tidak terganggu oleh aktivitas "kurang kritis" seperti streaming Spotify di bagian admin atau browsing berat di VLAN guest. Tapi, bagaimana caranya? Apakah kita langsung terjun ke dunia gelap mangle dan queue tree yang kita bahas sebelumnya?
Tidak. Kerja sunyi yang pertama adalah memetakan. Sebelum mengetik satu pun perintah di router, saya buat spreadsheet. Satu kolom untuk VLAN, satu kolom untuk perkiraan jumlah pengguna aktif, satu kolom untuk aplikasi kritis yang digunakan, dan satu kolom untuk "kebutuhan bandwidth wajar".
Proses menentukan "kebutuhan wajar" ini adalah seni negosiasi dengan realitas. Kamu tidak bisa datang ke kepala radiologi dan bilang, "Bapak, saya kasih bandwidth 2Mbps saja ya, cukup untuk lihat gambar rontgen." Kamu akan diusir. Tapi kamu juga tidak bisa memberi 50Mbps untuk 3 komputer saja. Jadi, metode saya: observasi puncak. Saya pasang monitor traffic sederhana di gateway selama seminggu, lihat pola puncak penggunaan tiap VLAN. Dari situ, dapatlah angka kasar:
- Radiologi & Lab: Puncak bisa 30-40Mbps saat transfer gambar/data batch.
- Rekam Medis & Rawat Inap: Konstan 10-15Mbps untuk akses EMR.
- Manajemen & Admin: 5-10Mbps, banyak browsing dan email.
- Guest: Suka meledak, bisa 20Mbps kalau tidak dikendalikan.
Total *demand* potensial jauh melebihi 100Mbps link primary. Inilah masalah klasik. Sumber terbatas, keinginan tak terbatas.
Dengan arsitektur active-standby, solusi kompleks seperti mangle untuk menandai koneksi per ISP menjadi berlebihan. Kenapa? Karena semua traffic untuk internet, pada kondisi normal, hanya akan keluar melalui satu pintu saja: ether1 (ISP Primer). Pintu kedua (ISP Backup) benar-benar idle. Jadi, kita tidak perlu repot membuat classifier yang membagi koneksi ke dua ISP. Kita hanya perlu membagi bandwidth dari satu *parent* yang jelas: antarmuka ISP Primer.
Di sinilah pilihan jatuh pada **Simple Queue**, bukan Queue Tree yang lebih granular. Tunggu, bukankah tadi di bab sebelumnya Queue Tree disebut lebih powerful? Iya, untuk skenario load-balance atau multi-WAN kompleks. Tapi untuk kasus ini, Simple Queue memiliki keunggulan besar: kemudahan administrasi.
Dengan 13 VLAN, saya harus membuat dan memelihara aturan. Simple Queue di MikroTik punya fitur *Queue Types* yang sangat membantu. Saya bisa mendefinisikan sebuah tipe antrian bernama "HS-Corp-100M" dengan parameter pcq-rate=5M dan pcq-limit=50 untuk pcq-classifier=dst-address. Artinya, saya membuat sebuah templat yang berkata: "Di dalam VLAN ini, setiap *tujuan IP* (dst-address) yang berbeda akan mendapat jatah maksimal 5Mbps, dengan total batas sesuai yang kita set di antrian induknya."
Lalu, untuk setiap VLAN, saya buat satu Simple Queue. Misal untuk VLAN Radiologi (10.30.0.0/24):
/queue simple add name=VLAN30-Radiologi target=10.30.0.0/24 parent=ether1 queue=HS-Corp-100M/HS-Corp-100M max-limit=40M/40M
Membacanya: Buat antrian sederhana bernama VLAN30-Radiologi yang berlaku untuk subnet 10.30.0.0/24, turunan dari interface ether1 (ISP Primer), menggunakan templat tipe antrian "HS-Corp-100M" untuk upload dan download, dengan batas maksimal total 40Mbps untuk upload dan 40Mbps untuk download dari/ke internet.
Ajaibnya, dengan templat pcq-rate tadi, meski total VLAN itu diberi cap 40Mbps, tidak ada satu *host* pun di dalamnya yang bisa mendominasi sepenuhnya. Jika ada satu komputer mulai download besar dari satu server di internet, dia akan dibatasi max 5Mbps (sesuai pcq-rate), membiarkan bandwidth tersisa untuk komputer lain di VLAN yang sama yang butuh akses. Ini adalah keadilan tingkat mikro, dikelola otomatis oleh router.
Proses ini diulang untuk 12 VLAN lainnya. Farmasi dapat 15Mbps, Lab dapat 35Mbps, Guest dapat 10Mbps, dan seterusnya. Total yang dialokasikan sengaja melebihi 100Mbps (misal jadi 120Mbps). Itu bukan masalah, karena yang penting adalah *max-limit* per VLAN, bukan jaminan bandwidth. Ini prinsip *overcommit* yang umum di dunia jaringan. Tidak semua VLAN akan mencapai puncaknya secara bersamaan. Saat Radiologi sibuk pagi hari, Farmasi mungkin sedang sepi. Sistem akan membagi secara dinamis sesuai permintaan, asalkan tidak melampaui batas maksimal VLAN-nya.
Lalu, bagaimana dengan ISP Backup? Untuk itu, saya buat set aturan Simple Queue kedua yang identik, tapi dengan parent=ether2 dan max-limit yang disesuaikan dengan kapasitas backup 50Mbps. Lalu, aturan ini saya disable. Mengapa? Karena aturan Simple Queue hanya aktif jika parent-nya digunakan. Saat ISP Primer aktif, semua traffic pakai aturan queue dengan parent=ether1. Begitu failover terjadi dan route keluar beralih ke ether2, secara otomatis aturan queue dengan parent=ether2 akan menjadi aktif dan membatasi traffic sesuai template baru yang lebih kecil. Tidak perlu script, tidak perlu trigger. Itu kerja sunyi sistem yang elegan.
Tantangan terbesarnya justru non-teknis: sosialisasi. Setelah aturan berlaku, pasti ada yang protes. "Kok buka YouTube di lobby jadi lambat?" Ya, karena kamu ada di VLAN Guest yang saya batasi 10Mbps dengan pcq-rate 1Mbps per host. "Transfer data dari mesin Lab ke cloud jadi lama!" Itu mungkin karena saat itu sedang mencapai batas 35Mbps untuk VLAN Lab, dan pcq-rate 5Mbps membatasi satu sesi transfer besar. Solusinya? Edukasi, dan terkadang, penyesuaian fine-tuning. Mungkin pcq-rate untuk VLAN Lab perlu dinaikkan jadi 10Mbps karena memang ada kebutuhan transfer file besar dari satu host. Itu bisa diubah dalam hitungan menit.
Setelah sebulan berjalan, grafiknya jadi cerita yang menarik. Garis untuk VLAN Guest yang dulu suka melonjak tinggi, sekarang terjebak di plafon 10Mbps. Garis untuk VLAN Lab dan Radiologi naik-turun sesuai jam sibuk, tapi tidak pernah saling menjegal karena masing-masing punya "jalur" sendiri yang aman. Yang paling puas adalah ketika ISP Primer down tengah malam, dan failover ke backup terjadi. Grafik langsung "turun kelas", semua antrian mengikuti aturan parent ether2 yang lebih ketat, tapi layanan kritis tetap berjalan—hanya lebih lambat. Tidak ada yang crash. Tidak ada yang panik. Itulah tujuan sebenarnya: predictability, predictability, predictability. Bukan kecepatan maksimal, tapi perilaku jaringan yang bisa diprediksi dan dikelola.
Kesimpulan dari studi kasus ini? Tool yang canggih (Queue Tree/Mangle) tidak selalu jadi pilihan terbaik. Pemahaman mendalam terhadap arsitektur aktual (single active gateway) dan kebutuhan operasional (kemudahan maintenance) bisa mengarahkan kita pada solusi yang lebih sederhana (Simple Queue dengan PCQ). Kekuatan solusi ini justru terletak pada kesederhanaannya. Siapa pun admin yang menggantikan saya kelak, bisa membuka konfigurasi dan dalam 5 menit paham: "Oh, ada 13 antrian, masing-masing untuk VLAN, pakai templat yang sama, parent-nya ether1." Bandingkan jika dia harus melacak aturan mangle yang menandai koneksi, lalu queue tree yang merujuk pada packet-mark, di tengah puluhan aturan firewall lainnya. Itu bisa jadi mimpi buruk.
Dalam kerja sunyi sistem, pemahaman mendalam tentang arsitektur yang ada seringkali mengarah pada solusi yang terlihat sederhana di permukaan, namun justru itulah yang paling robust dan mudah dikelola di balik layar.
Tanya Jawab Seputar Implementasi
Q: Kenapa pakai PCQ (pcq-rate), bukan langsung bagi rata max-limit saja?
A: Karena max-limit hanya membatasi total VLAN. Tanpa PCQ, satu orang di dalam VLAN bisa monopoli bandwidth itu dan membuat yang lain kelaparan. PCQ memastikan pembagian yang adil di level pengguna/device, secara otomatis.
Q: Apakah tidak repot maintain 13 aturan queue? Bagaimana kalau ada penambahan VLAN?
A> Justru dengan Simple Queue, penambahan VLAN sangat mudah. Cukup duplikat satu queue yang sudah ada, ganti `target` subnet-nya, sesuaikan `max-limit`-nya. Lebih cepat daripada harus setup mangle rules baru dan menyesuaikan queue tree.
Q: Bagaimana memonitor apakah aturan ini efektif?
A> Gunakan grafik di menu Queue. Lihat, apakah ada VLAN yang konsisten mencapai `max-limit`-nya? Jika iya, mungkin perlu review kebutuhan. Juga, pantau keluhan. Jika keluhan "lambat" menjadi spesifik ke satu VLAN, artinya isolasi masalah berhasil.
Q: Bagaimana dengan lalu lintas internal antar VLAN (server farm ke radiologi)?
A> Simple Queue dengan `parent=ether1` hanya membatasi traffic yang keluar melalui interface tersebut (ke internet). Lalu lintas antar VLAN yang melalui internal bridge/router tidak akan kena limit. Ini sesuai desain.
Q: Apa yang terjadi jika ISP Primary kembali hidup setelah failover?
A> Routing akan kembali ke ether1. Karena aturan queue untuk parent=ether1 dalam status enabled, maka aturan bandwidth 100Mbps akan otomatis berlaku kembali. Traffic akan beralih tanpa intervensi.
Q: Apakah skema overcommit (total limit > 100Mbps) berisiko?
A> Berisiko hanya jika semua VLAN benar-benar mencapai puncak bersamaan, yang jarang terjadi. Namun, risiko itu adalah trade-off untuk fleksibilitas. Jika terjadi, bandwidth akan dibagi secara kontesional sesuai max-limit masing-masing. Yang penting, layanan kritis tetap punya plafon yang terjamin (misal, Radiologi 40Mbps), sehingga tidak akan terdesak sepenuhnya oleh traffic lain.
Q: Apa pelajaran terbesar dari studi kasus ini?
A> Jangan jadi "kutu buku teknis" yang memaksakan solusi paling kompleks. Dengarkan jaringanmu. Pahami polanya. Seringkali, solusi elegan itu adalah yang paling sesuai dengan pola yang sudah ada, bukan yang paling indah di teori.
Bandwidth Management in a Hospital Network: A Case Study of 13 VLANs and Dual ISP
After understanding the principles of granularity using mangle and queue tree, let's see how those principles are applied—or even simplified—in an environment with real needs but a relatively straightforward routing architecture.
So, here's the story. There's a medium-sized private hospital on the outskirts of town. They have 13 VLANs. No small number. These VLANs represent their own worlds: there's VLAN for Management, IT, Finance, HRD, Medical Records, Radiology, Laboratory, Pharmacy, Outpatient, Inpatient, Nutrition, Engineering, and of course—Guest. Each department has its own microcosm, with dedicated servers, dedicated printers, and of course, dedicated internet complaints.
Internet connection is dual ISP. One 100Mbps fiber optic from a major provider as *primary*. The other, a 50Mbps radio link from a local provider as *backup*. The configuration is pure active-standby. No complex load-balancing. If the fiber fails, all traffic is routed to the radio. Simple, predictable. This is a crucial point. Because the routing is simple, the logic for bandwidth management can be simple too. This is our first lesson: don't force a complex solution if the base architecture is already simple.
The main requirement is clear: ensure critical hospital traffic—such as access to lab applications, digital radiology (PACS), and electronic medical records—is not disrupted by "less critical" activities like Spotify streaming in the admin section or heavy browsing on the Guest VLAN. But, how? Do we immediately dive into the dark world of mangle and queue tree we discussed earlier?
No. The first quiet work is mapping. Before typing a single command on the router, I created a spreadsheet. One column for VLAN, one for estimated number of active users, one for critical applications used, and one for "reasonable bandwidth need."
The process of determining "reasonable need" is an art of negotiation with reality. You can't go to the head of radiology and say, "Sir, I'll give you 2Mbps bandwidth, enough to view x-rays." You'll be thrown out. But you also can't give 50Mbps to just 3 computers. So, my method: peak observation. I installed simple traffic monitoring on the gateway for a week, observed the peak usage pattern of each VLAN. From there, rough numbers emerged:
- Radiology & Lab: Peak could be 30-40Mbps during batch image/data transfer.
- Medical Records & Inpatient: Constant 10-15Mbps for EMR access.
- Management & Admin: 5-10Mbps, lots of browsing and email.
- Guest: Prone to spikes, could hit 20Mbps if uncontrolled.
The total potential *demand* far exceeds the 100Mbps primary link. This is the classic problem. Limited supply, unlimited wants.
With an active-standby architecture, complex solutions like mangle to mark connections per ISP become overkill. Why? Because all internet traffic, under normal conditions, will only exit through one door: ether1 (Primary ISP). The second door (Backup ISP) is completely idle. So, we don't need to bother creating a classifier that splits connections to two ISPs. We only need to divide the bandwidth from one clear *parent*: the Primary ISP interface.
This is where the choice fell on **Simple Queue**, not the more granular Queue Tree. Wait, wasn't Queue Tree said to be more powerful in the previous chapter? Yes, for complex load-balance or multi-WAN scenarios. But for this case, Simple Queue has a major advantage: ease of administration.
With 13 VLANs, I have to create and maintain rules. Simple Queue in MikroTik has a very helpful *Queue Types* feature. I can define a queue type named "HS-Corp-100M" with parameters pcq-rate=5M and pcq-limit=50 for pcq-classifier=dst-address. This means, I'm creating a template that says: "Within this VLAN, each different *destination IP* (dst-address) will get a maximum quota of 5Mbps, with the total limit as we set in the parent queue."
Then, for each VLAN, I create one Simple Queue. For example, for the Radiology VLAN (10.30.0.0/24):
/queue simple add name=VLAN30-Radiology target=10.30.0.0/24 parent=ether1 queue=HS-Corp-100M/HS-Corp-100M max-limit=40M/40M
Reading it: Create a simple queue named VLAN30-Radiology that applies to subnet 10.30.0.0/24, a child of interface ether1 (Primary ISP), using the queue type template "HS-Corp-100M" for both upload and download, with a total maximum limit of 40Mbps for upload and 40Mbps for download to/from the internet.
The magic is, with that pcq-rate template, even though the total VLAN is capped at 40Mbps, no single *host* within it can completely dominate. If one computer starts a large download from a server on the internet, it will be limited to a max of 5Mbps (according to pcq-rate), leaving the remaining bandwidth for other computers in the same VLAN that need access. This is micro-level fairness, managed automatically by the router.
This process is repeated for the other 12 VLANs. Pharmacy gets 15Mbps, Lab gets 35Mbps, Guest gets 10Mbps, and so on. The total allocated intentionally exceeds 100Mbps (e.g., becomes 120Mbps). That's not a problem, because what matters is the *max-limit* per VLAN, not guaranteed bandwidth. This is the principle of *overcommit* common in networking. Not all VLANs will peak simultaneously. When Radiology is busy in the morning, Pharmacy might be idle. The system will dynamically share according to demand, as long as it doesn't exceed each VLAN's maximum limit.
What about the Backup ISP? For that, I create a second, identical set of Simple Queue rules, but with parent=ether2 and max-limit adjusted to the 50Mbps backup capacity. Then, I disable these rules. Why? Because Simple Queue rules are only active if their parent is being used. When the Primary ISP is active, all traffic uses the queue rules with parent=ether1. Once a failover occurs and the exit route switches to ether2, the queue rules with parent=ether2 will automatically become active and limit traffic according to the new, smaller template. No scripts needed, no triggers. That's the elegant, quiet work of the system.
The biggest challenge was actually non-technical: socialization. Once the rules are applied, there will be complaints. "Why is YouTube in the lobby so slow?" Well, because you're on the Guest VLAN which I limited to 10Mbps with a pcq-rate of 1Mbps per host. "Data transfer from the Lab machine to the cloud is now slow!" That might be because the VLAN Lab was hitting its 35Mbps cap at that time, and the pcq-rate of 5Mbps was limiting one large transfer session. The solution? Education, and sometimes, fine-tuning adjustments. Maybe the pcq-rate for the Lab VLAN needs to be increased to 10Mbps because there is indeed a need for large file transfers from a single host. That can be changed in minutes.
After a month of operation, the graph tells an interesting story. The line for the Guest VLAN, which used to spike high, is now stuck at a 10Mbps ceiling. The lines for Lab and Radiology VLANs go up and down according to busy hours, but never trip over each other because each has its own safe "lane." The most satisfying moment was when the Primary ISP went down in the middle of the night, and failover to the backup occurred. The graph immediately "downgraded," all queues followed the stricter parent=ether2 rules, but critical services kept running—just slower. Nothing crashed. No one panicked. That's the real goal: predictability, predictability, predictability. Not maximum speed, but network behavior that is predictable and manageable.
The conclusion from this case study? Fancy tools (Queue Tree/Mangle) are not always the best choice. A deep understanding of the actual architecture (single active gateway) and operational needs (ease of maintenance) can lead us to a simpler solution (Simple Queue with PCQ). The strength of this solution lies precisely in its simplicity. Any admin who replaces me in the future can open the configuration and understand in 5 minutes: "Oh, there are 13 queues, each for a VLAN, using the same template, parent is ether1." Compare that if they had to trace mangle rules marking connections, then queue trees referencing packet-marks, amidst dozens of other firewall rules. That could be a nightmare.
In the quiet work of the system, a deep understanding of the existing architecture often leads to solutions that look simple on the surface, yet that is precisely what makes them the most robust and easy to manage behind the screen.
Q&A on Implementation
Q: Why use PCQ (pcq-rate), not just split the max-limit evenly?
A: Because max-limit only limits the total VLAN. Without PCQ, one person inside the VLAN could monopolize that bandwidth and starve others. PCQ ensures fair sharing at the user/device level, automatically.
Q: Isn't it cumbersome to maintain 13 queue rules? What if a VLAN is added?
A> On the contrary, with Simple Queue, adding a VLAN is very easy. Just duplicate an existing queue, change the `target` subnet, adjust its `max-limit`. Much faster than setting up new mangle rules and adjusting queue trees.
Q: How to monitor if these rules are effective?
A> Use the graph in the Queue menu. See, is any VLAN consistently hitting its `max-limit`? If yes, its needs might need reviewing. Also, monitor complaints. If "slow" complaints become specific to one VLAN, it means problem isolation is working.
Q: What about internal traffic between VLANs (server farm to radiology)?
A> Simple Queue with `parent=ether1` only limits traffic exiting through that interface (to the internet). Traffic between VLANs passing through the internal bridge/router will not be limited. This is by design.
Q: What happens when the Primary ISP comes back online after a failover?
A> Routing will revert to ether1. Since the queue rules for parent=ether1 are enabled, the 100Mbps bandwidth rules will automatically apply again. Traffic will switch over without intervention.
Q: Isn't the overcommit scheme (total limit > 100Mbps) risky?
A> It's only risky if all VLANs truly peak simultaneously, which is rare. However, that risk is the trade-off for flexibility. If it happens, bandwidth will be contested according to each VLAN's max-limit. The important thing is that critical services still have a guaranteed ceiling (e.g., Radiology 40Mbps), so they won't be completely squeezed out by other traffic.
Q: What's the biggest lesson from this case study?
A> Don't be a "technical bookworm" who forces the most complex solution. Listen to your network. Understand its patterns. Often, the elegant solution is the one that best fits the existing pattern, not the one that looks most beautiful in theory.
Thank you for stopping by! If you enjoy the content and would like to show your support, how about treating me to a cup of coffee? �� It’s a small gesture that helps keep me motivated to continue creating awesome content. No pressure, but your coffee would definitely make my day a little brighter. ☕️ Buy Me Coffee

Post a Comment for "Manajemen Bandwidth di Jaringan Rumah Sakit: Studi Kasus 13 VLAN dan Dual ISP"
Post a Comment
You are welcome to share your ideas with us in comments!