Flowbix · AWS ELB · Balanceadores
TODO TRÁFEGO BALANCEADO, UMA SÓ TELA.
Monitoramento de Application e Network Load Balancers em tempo real — requests e conexões, latência p50/p90/p99, códigos HTTP 2xx–5xx, target groups saudáveis, flows TCP/TLS/UDP e resets. Do ALB ao NLB, consolidado numa única sala de operação sobre Zabbix (CloudWatch).
99,98%
Respostas 2xx · balanceadores
/ BALANCEADORES
ATIVOS
18
ALB12
NLB6
/ REQUESTS/MIN
AO VIVO
Requests24,8k/min
LCUs42
/ LATÊNCIA · p99
OK
142 ms
p50 · p9038 · 96 ms
Target resp.ok
/ HTTP 5xx
ATENÇÃO
12
4xx84
TLS errors2
/ TARGET GROUPS
SAUDÁVEL
236/240
Grupos32
Unhealthy4
/ NLB FLOWS
AO VIVO
48 k
Resets3
PPS pico1,2 M
Sala de operação de rede
Balanceadores ELB em tempo real
18
Balanceadores
99,98%
Respostas 2xx
142ms
Latência p99
236/240
Hosts OK
24×7
Monitoramento
01/08
Frota de balanceadoresCloudWatch → Zabbix → NOC
AWS/ApplicationELB12 ALB
AWS/NetworkELB6 NLB
LLD · Target Groups40 TG
LLD · CloudWatch Alarms46
poll 60 s
NOC Cloud Flowbix
ALB · NLB · Target Groups · Alarmes — uma tela, ao vivo
Índice operacional consolidado97,2/100
18/18
Balanceadores · 12 ALB · 6 NLB
24,8k/min
Requests · pico 31,2k
142ms
Latência p99 · p50 24 ms
12/min
HTTP 5xx · taxa 0,05%
236/240
Targets healthy · 4 unhealthy
2 ALARM
Alarmes · 44 OK · 0 insuf.
01 · RequestsAO VIVO
Visão & Requests
Requests, conexões, LCUs e throughput da frota ALB.
Requests24,8k/m
Conex. ativas18,4k
Novas conex.2,1k/s
Rejeitadas0
Bytes proc.3,2 GB/m
Regras aval.41k/m
Redirects1,2k/m
LCUs ALB148
Requests · k/min · 24 hmáx 31,2
−24 h−12 hagora
Requests por ALB · top 3alb-web
24,8k req/min · 12 ALB · 3,2 GB/min
02 · ALB HTTPp99 ALTA
ALB Latência & HTTP
Target response time p50/p90/p99 e códigos HTTP.
p5024 ms
p9078 ms
p99142 ms
HTTP 2xx98,7%
HTTP 4xx312/m
HTTP 5xx12/m
ELB 5xx3
TLS neg. err0
p99 · ms · 24 hmáx 210
−24 h−12 hagora
Erros 5xx · bins 1,5 h3 janelas
p99 142 ms · 5xx 12/min · gatilho 5xx>5
03 · Targets236/240
Target Groups & Saúde
Healthy/unhealthy hosts, roteamento e por-target.
Target groups40
Healthy236
Unhealthy4
Req/target620/m
Anômalos1
Mitigados0
Estado rotaOK
Health DNSOK
Healthy hosts · 24 hmín 232
−24 h−12 hagora
Saúde por TG · 16 grupos4 unh.
236/240 healthy · 4 unhealthy · gatilho >0
04 · NLBEM REDE
NLB Flows & Resets
Flows TCP/UDP/TLS, resets, pacotes e alocação de porta.
Flows ativos42k
Novos flows3,4k/s
Pico pkts/s1,2M
Reset client22/m
Reset elb18/m
Reset target4/m
Port alloc err0
LCUs NLB62
Flows ativos · k · 24 hmáx 51
−24 h−12 hagora
TCP resets · por origemelb 18
42k flows · 6 NLB · 0 port errors
05 · Alarmes2 ALARM
Alarmes & Capacidade
Alarmes CloudWatch, LCUs consumidas e throughput.
Alarmes46
Estado OK44
Em ALARM2
Insuf. dados0
LCUs total210
LCU ALB148
LCU NLB62
Bytes/dia4,6 TB
LCUs consumidas · 24 hmáx 248
−24 h−12 hagora
Capacidade consumida210 LCU
2 em ALARM · 44 OK · 210 LCUs
Flowbix · AWS ELB · 5 dashboards numa sessão
02 / 08
Requests6 ALB · agregado
24,8k/min
RequestCount · pico 24h 31,2k/min · 1,49M/h
Active conn.estável
3.180
ActiveConnectionCount · pico 4.120
New conn.reuse 96%
1,26k/min
NewConnectionCount · média 1,10k
Rejectedacima da base
18/min
RejectedConnectionCount · base <5 · surge queue
Consumed LCUscusto · billing
42 LCU
Σ 6 ALB · processados 91,3 GB/h · pico 58 LCU
Requests & conexões · 24h1 min · média
Requests/minActive connections- - limite 30k/min
32,0 k24,0 k16,0 k8,0 k0,0 k
00h04h08h12h16h20h24h
Requests/min (área âmbar) · Active conn. (ciano · escala 0–4,5k)pico 31,2k · 4.120 conn · limite 30k
Requests por load balancerΣ 24,8k/min
Processed bytes · agregado91,3 GB/h
Rule evaluations148k/min
Tamanho médio de request14,2 KB
Load balancers (ALB)6 ativos · sa-east-1
| Load balancer | Req/min | Active | New/min | LCUs | Bytes/h | Status |
|---|---|---|---|---|---|---|
| alb-web-prod | 9.850 | 1.180 | 460 | 16,4 | 42,8 GB | Normal |
| alb-api-prod | 7.420 | 940 | 380 | 12,8 | 28,1 GB | Normal |
| alb-checkout-prod | 3.180 | 520 | 210 | 6,9 | 9,4 GB | 5xx subindo |
| alb-static-01 | 1.910 | 190 | 78 | 2,1 | 6,7 GB | Normal |
| alb-admin-01 | 1.640 | 240 | 92 | 2,8 | 3,2 GB | Normal |
| alb-partners-01 | 800 | 110 | 44 | 1,0 | 1,1 GB | Normal |
| Σ · 6 ALB | 24.800 | 3.180 | 1.264 | 42,0 | 91,3 GB | 6/6 up |
Conexões & throughputkeep-alive alto
Active connections3.180
New connections1,26k/min
Rejected connections18/min
Processed bytes91,3 GB/h
HTTP redirects2,14k/min
HTTP fixed response640/min
Flowbix · ALB · requests, conexões, LCUs e throughput
03 / 08
Target resp. p99SLA 150 ms
142ms
meta <150 ms · pico 24h 162
Target resp. p50saudável
38ms
mediana · -4 ms vs 24h
Target resp. p90estável
96ms
p90 · pico 24h 118 ms
HTTP 5xxgatilho >5
12
ELB 5xx · últ. 5 min · 3 ALBs
HTTP 4xxgatilho >5
84
últ. 5 min · 401 22 · 403 9 · 404 41 · 429 12
Target response time · 24h (p50 / p90 / p99)1 pico acima do SLA
p5038 ms
p9096 ms
p99142 ms
máx p99 24h162 ms
acima SLA 150 ms1 pico
200 ms150 ms100 ms50 ms0 ms
00h04h08h12h16h20h24h
p50p90p99SLA 150 mstarget response time · ms
Códigos HTTP · 5 min1,90 M req
5xx por código · gerados pelo ELB
500 · Internal Server Error3
502 · Bad Gateway5
503 · Service Unavailable3
504 · Gateway Timeout1
HTTP por load balancer5 ALB · 24h
| ALB | 2xx | 3xx | 4xx | 5xx | p99 | TLS err |
|---|---|---|---|---|---|---|
| alb-web-prod | 1,24 M | 28,4 k | 61 | 7 | 142 ms | 2 |
| alb-api-prod | 402 k | 9,1 k | 18 | 4 | 96 ms | 0 |
| alb-checkout-prod | 96,2 k | 2,3 k | 4 | 1 | 168 ms | 0 |
| alb-admin-int | 41,8 k | 1,1 k | 1 | 0 | 74 ms | 0 |
| alb-media-cdn | 58,6 k | 380 | 0 | 0 | 52 ms | 0 |
Consolidado · 4xx / 5xx (5 min)84 · 12
p99 mais alto · alb-checkout-prod168 ms
Erros & autenticação2 gatilhos
HTTP 5xx (ELB)12 · >5
Client TLS negotiation err3 /min
Target TLS negotiation err1 /min
Target connection errors6 /5min
Target reset count (RST)24
Auth error / falha2 / 0
Auth latency · p9588 ms
Auth success rate99,4%
Flowbix · ALB · latência p50/p90/p99 e códigos HTTP
04 / 08
Target groupsALB+NLB
32
5 balanceadores · descoberta LLD
Hosts saudáveis98,3%
236/240
27 tg em 100% · 4 hosts fora de rotação
Unhealthy3 tg
4
trigger unhealthy>0
AnomalousML
2
1 auto-mitigado
Request/target médioRPS/target
128/min
pico 214/min · tg-img-cdn
Target groups · saúde (32)3 degradados
100% saudável
draining / registrando
degradado
crítico
241812812976
5411610895
146856746
4581510969
27 saudáveis · 2 draining · 2 degradados · 1 crítico1 bloco = 1 tg · nº = hosts healthy
Healthy vs unhealthy hosts · 24h2 séries · eixo duplo
240237234231228
129630
00h04h08h12h16h20h24h
healthy (esq · hosts) unhealthy (dir · hosts)- - meta healthy ≥ 236 · pico unhealthy 6 · 12h40
Target groups8 de 32 · ordenado por risco
| Grupo | Load balancer | Healthy | Unhealthy | Req/target | 5xx target | Estado |
|---|---|---|---|---|---|---|
| tg-grpc-01 | nlb-ingest-01 | 6/8 | 2 | 54/min | 1,12% | Crítico |
| tg-checkout | alb-web-prod | 11/12 | 1 | 96/min | 0,41% | Degradado |
| tg-img-cdn | alb-edge-01 | 15/16 | 1 | 214/min | 0,08% | Degradado |
| tg-api-01 | alb-web-prod | 24/24 | 0 | 182/min | 0,02% | Normal |
| tg-web-01 | alb-web-prod | 18/18 | 0 | 143/min | 0,00% | Normal |
| tg-search-01 | alb-web-prod | 12/12 | 0 | 118/min | 0,03% | Normal |
| tg-auth-01 | alb-edge-01 | 8/8 | 0 | 71/min | 0,01% | Normal |
| tg-batch-01 | alb-internal | 9/9 | 0 | 22/min | 0,00% | Normal |
+24 grupos · 133 hosts · 100% saudáveisΣ 32 tg · 240 hosts · 236 healthy · 5xx global 0,06%
Saúde de roteamento2 anomalias
Anomalous hosts (deteção ML)2
Mitigated hosts (auto-mitigação)1
Unhealthy routing requests · 5m312
Target connection errors · 24h41 · 0,03%
Client TLS negotiation errors3
Reset TCP (NLB · target) · min128
Flowbix · Target groups · healthy/unhealthy hosts por grupo
05 / 08
Active flows6 NLBs
48,2 k
pico 24h 48,3k · cap. est. 90k
New/s1,9k
Encerr./s1,7k
TCP40,1k
TLS+UDP8,1k
ActiveFlowCount agregado · TCP 83% · TLS 13% · UDP 4%
New flows/s/s
1,9 k
pico 4,2k/s
1 min1,84k
5 min1,71k
NewFlowCount · sem surtos
TCP resetsestável
3 /s
elb 0 · target 3
cliente (normal)41/s
RST 5 min63
TCP_Reset_Count · elb/target OK
Peak packets/s24 h
1,2 M
agora 1,18M pps · média 0,74M
Pacotes/s1,18M
Bytes/s1,78 GB
Jumbo8,4%
Frag.0,3%
PeakPacketsPerSecond · sem port-exhaustion
Consumed LCUscusto
42,8
TCP 31,4 · TLS 9,1
UDP LCU2,3
Custo/diaR$ 148
ConsumedLCUs · dim. dominante: flows
Active flows · 24h (TCP / TLS / UDP)3 protocolos · ActiveFlowCount
50 k38 k25 k13 k0
00h04h08h12h16h20h24h
TCPTLSUDPcap. 48k
min 31,4k · méd 41,2k · máx 48,3k · agora 48,2k
Resets & errossob controle
Port allocation errors0
SG blocked flows · entrada1.204
SG blocked flows · saída86
Unhealthy routing flows27
Client/target TLS neg. errors4
PortAllocationError 0 → SNAT saudável · gatilho unhealthy>0 em nlb-voip-01
NLB · por balanceador6 NLB · 2 AZ · cross-zone on
| NLB | Active | New/s | Bytes 24h | PPS | RST 5min | Unhealthy rout. |
|---|---|---|---|---|---|---|
| nlb-ingest-01 | 21,4k | 820 | 3,8 TB | 480k | 12 | 0 |
| nlb-ingest-02 | 14,8k | 610 | 2,4 TB | 310k | 8 | 0 |
| nlb-rtmp-01 | 6,2k | 180 | 5,1 TB | 210k | 40 | 6 |
| nlb-mqtt-01 | 3,1k | 240 | 0,6 TB | 88k | 2 | 0 |
| nlb-syslog-01 | 1,9k | 90 | 0,9 TB | 61k | 0 | 0 |
| nlb-voip-01 | 0,8k | 45 | 0,3 TB | 52k | 1 | 21 |
Σ 48,2k flows · 1,9k new/s · 13,1 TB · 1,20M pps · 63 RST · 1 NLB c/ rota unhealthy (voip-01 · gatilho)
Fluxos por protocolodistribuição · 24h
Bytes 24h · TCP / TLS / UDP9,2 / 3,1 / 0,7 TB
LCU · TCP / TLS / UDP31,4 / 9,1 / 2,3
Peak pps · TCP / TLS / UDP0,96M / 0,18M / 0,06M
SG blocked · entrada1.204
SG blocked · saída86
ProcessedBytes/Packets por listener · Σ LCU 42,8
Flowbix · NLB · flows TCP/TLS/UDP, resets e pps
06 / 08
AlarmesCloudWatch LLD
42 alarmes
37 OK · 2 s/ dados · 3 ALARM
Em ALARMação
3 ativos
5xx · latency · unhealthy · 2 un-ack
LCUs total · ALB+NLBagora
214 LCU
pico 24h 218 · orçam. 260 · headroom 18%
LCUs ALB5 ALB
138 LCU
64% do total · pico 141 · US$ 0,008/LCU-h
LCUs NLB4 NLB
76 LCU
36% · pico 76
Consumed LCUs · 24h (ALB + NLB)Σ 214 LCU · pico 218
280 LCU22016010040 LCU
00h04h08h12h16h20h24h
ALBNLB· orçamento total 260 LCU
agora 214 · - - alarme LCU ALB > 160
Consumed LCUs por balanceador5 ALB · 4 NLB
Alarmes ativos · CloudWatch3 ALARM · 2 s/ dados · aval. 60 s
| Alarme | Balanceador | Métrica | Estado | Razão | Idade |
|---|---|---|---|---|---|
| alb-checkout-5xx-high | alb-checkout-prod ALB | HTTPCode_ELB_5XX | ALARM | 5xx 11/min > 5 | 6 min |
| tg-api-01-unhealthy | alb-api-prod ALB | UnHealthyHostCount | ALARM | 1 de 4 alvos DOWN | 14 min |
| alb-api-p99-latency | alb-api-prod ALB | TargetResponseTime | ALARM | p99 1,82 s > 1,50 | 22 min |
| nlb-ingest-01-hosts | nlb-ingest-01 NLB | UnHealthyHostCount | S/ DADOS | sem datapoints (3×) | 9 min |
| alb-web-4xx-high | alb-web-prod ALB | HTTPCode_ELB_4XX | S/ DADOS | coleta intermitente | 4 min |
| nlb-ingest-02-flows | nlb-ingest-02 NLB | ActiveFlowCount | OK | normalizado 11:41 | 2 min |
Capacidade & coletaLCU / CloudWatch
Proxy de custo · LCU-horaUS$ 37/dia · ~1,1k/mês
Health checks · alvos512/min · 99,2% ok
API CloudWatch · GetMetricData480 ms · 0 throttle
Cobertura · LB / target groups9 LB · 46 TG · 60 s
Flowbix · Alarmes CloudWatch, LCUs e capacidade
07 / 08
Próximo passo
Vamos centralizar?
Agende uma demo da Missão Crítica Flowbix, fale com o time comercial e traga energia, no-break, gerador, CFTV e controle de acesso dos seus sites críticos para uma única sala de operação.
service
FLOWBIX
Comercial · Flowbix ✓
Conta verificada
Escaneie para iniciar uma conversa
Flowbix
·
Missão Crítica · Monitoramento de Sites Críticos
08 / 08