Example queries for common monitoring tasks - Amazon ElastiCache
Services or capabilities described in AWS documentation might vary by Region. To see the differences applicable to the AWS European Sovereign Cloud Region, see the AWS European Sovereign Cloud User Guide.

Example queries for common monitoring tasks

The following queries answer common questions about a node-based Valkey replication group. Each one is ready to copy into a CloudWatch dashboard, and you can edit any of them to suit your own workload.

Every query is grouped by topic and collapsed. Expand an entry to see the query it contains.

Some queries use metrics that are emitted only while they are selected in a resource metrics configuration. These queries are marked Requires detailed monitoring, and return no data until you select the metrics that they name. For more information, see Turning on detailed monitoring.

Note

These queries assume that you understand the PromQL constructs they use. For an introduction to each construct, with worked examples, see Querying your metrics.

Substituting your own values

Replace the following placeholders with values from your own replication group.

Placeholder Replace with
my-cluster Your replication group ID
my-cluster-0001-001 A node ID
prod-.* A regular expression that matches the replication group IDs you want
Note

A query can return at most 500 time series. Queries that break results down by node and by command can exceed that limit on a large replication group, in which case the response is truncated. Aggregate the result, or narrow it with a more specific label matcher.

Note

The ranges in these queries, such as [5m], suit a metric emitted every 60 seconds. If you have selected a metric, it is emitted every 15 seconds and you can shorten its ranges to [1m] for a more responsive result.

Turning a query into an alarm

A PromQL alarm carries its threshold inside the query. To turn any of these queries into an alarm, append a comparison. The query then returns a series only while the condition is true, and the alarm tracks each returned series separately. For example, the following query returns a series for each node that is using more than 90 percent of its memory limit.

100 * {"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"} / {"valkey.memory.max", "@resource.aws.elasticache.replication_group.id"="my-cluster"} > 90

Choose a threshold that suits your workload. A query that already ends in a comparison, such as == 0, can be used as an alarm unchanged. For how a PromQL alarm evaluates a query, see Creating alarms from queries.

Memory and capacity

100 * {"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"} / {"valkey.memory.max", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

A ratio against the configured limit, so it is comparable across different configurations.

100 * predict_linear({"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[1h], 3600) / on ("@resource.aws.elasticache.node.id") {"valkey.memory.max", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

predict_linear fits the trend of the last hour and projects it one hour ahead, so a value over 100 means that the node is projected to reach maxmemory within the hour. A linear fit does not model a workload that grows in steps.

100 * ( {"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"} - {"valkey.memory.not_counted_for_evict", "@resource.aws.elasticache.replication_group.id"="my-cluster"} ) / {"valkey.memory.max", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

The engine excludes some memory when deciding whether to evict, so this tracks the eviction decision more closely than total memory use does.

Requires detailed monitoring: valkey.memory.used.overhead

100 * {"valkey.memory.used.overhead", "@resource.aws.elasticache.replication_group.id"="my-cluster"} / {"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

A high share usually points to client buffer growth or replication backlog rather than dataset size.

max_over_time({"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[24h])

The range sets the period, up to 7 days.

deriv({"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[30m])

deriv is used because valkey.memory.used can decrease as well as increase, so rate does not apply.

CPU and compute

100 * sum without ("cpu.mode") ( rate({"process.cpu.time", "thread.type"="main", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

The main thread is single-threaded, so approaching 100 percent means the engine is saturated regardless of how many vCPUs the node has.

100 * sum by ("cpu.mode", "@resource.aws.elasticache.node.id") ( rate({"process.cpu.time", "thread.type"="main", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
deriv({"system.paging.usage", "system.paging.state"="used", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[30m])

A sustained positive value means the host is still moving pages to swap, which is worse than a steady non-zero figure.

rate({"process.paging.faults", "system.paging.fault.type"="major", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Only major faults are reported. They require disk access, so a sustained rate indicates memory pressure.

Throughput and commands

rate({"valkey.commands.processed", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])
sum by ("valkey.command.type") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
100 * sum(rate({"valkey.command.calls", "valkey.command.type"="write", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / sum(rate({"valkey.command.calls", "valkey.command.type"=~"read|write", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Commands that neither read nor write, such as PING, carry no valkey.command.type, so the denominator counts reads and writes only.

topk(10, sum by ("valkey.command") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])))
sum by ("valkey.command") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
sum by ("valkey.command") ( rate({"valkey.command.calls", "valkey.command"=~"GET|MGET|HGET|LRANGE", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
sum by ("valkey.command.category") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
sum(rate({"valkey.command.calls", "valkey.command.category"="sortedset", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
sum by ("@resource.aws.elasticache.shard.id") ( rate({"valkey.commands.processed", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Compare shards against each other to find imbalance.

rate({"valkey.commands.processed", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / on ("@resource.aws.elasticache.shard.id") group_left sum by ("@resource.aws.elasticache.shard.id") ( rate({"valkey.commands.processed", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Latency (per command)

1e6 * sum by ("valkey.command") ( rate({"valkey.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / (sum by ("valkey.command") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) > 0)

1e6 converts seconds to microseconds. The result is an average over the range.

topk(5, 1e6 * sum by ("valkey.command") (rate({"valkey.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / (sum by ("valkey.command") (rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) > 0))
1e6 * sum by ("valkey.command.type") ( rate({"valkey.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / (sum by ("valkey.command.type") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) > 0)
1e6 * sum by ("valkey.command.category") ( rate({"valkey.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / (sum by ("valkey.command.category") ( rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) > 0)
1e6 * sum(rate({"valkey.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / sum(rate({"valkey.command.calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Cache effectiveness

100 * sum(rate({"valkey.keyspace.hits", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) / (sum(rate({"valkey.keyspace.hits", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])) + sum(rate({"valkey.keyspace.misses","@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])))
rate({"valkey.keys.evicted", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Keys are evicted only under an eviction policy. Under noeviction, writes are rejected instead.

rate({"valkey.keys.expired", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])
rate({"valkey.keys.expired_fields", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Connections and clients

100 * {"valkey.clients.connected", "@resource.aws.elasticache.replication_group.id"="my-cluster"} / {"valkey.clients.max", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

A ratio against the configured limit, so it is comparable across different configurations.

rate({"valkey.connections.rejected", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])
rate({"valkey.connections.received", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

A high rate can indicate that clients are not pooling connections.

Client buffer pressure

Requires detailed monitoring: valkey.clients.output_buffer_disconnections

rate({"valkey.clients.output_buffer_disconnections", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

A disconnection is data the client never received.

Requires detailed monitoring: valkey.clients.query_buffer_disconnections

rate({"valkey.clients.query_buffer_disconnections", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

A disconnection is a request that was never completed.

Requires detailed monitoring: valkey.clients.evicted

rate({"valkey.clients.evicted", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Eviction means client memory reached its configured cap.

Replication and durability

{"valkey.replication.link_healthy", "valkey.role"="replica", "@resource.aws.elasticache.replication_group.id"="my-cluster"} == 0

Returns a series only for a replica whose link is down, so it can be used as an alarm unchanged.

Requires detailed monitoring: valkey.replication.sync.in_progress

{"valkey.replication.sync.in_progress", "@resource.aws.elasticache.replication_group.id"="my-cluster"} == 1
deriv({"valkey.replication.offset", "valkey.role"="primary", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Reported in bytes per second.

rate({"valkey.durability.buffer_exceeded_errors", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Every occurrence is a rejected write.

Persistence

{"valkey.persistence.rdb.save_in_progress", "@resource.aws.elasticache.replication_group.id"="my-cluster"} == 1

Expected during backups and diskless replication.

Network

100 * rate({"system.network.io", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / on ("@resource.aws.elasticache.node.id") group_left {"system.network.bandwidth.limit", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

A ratio against the limit the instance reports, so it is comparable across node types. The limit applies to each direction separately, so each direction is compared against the whole limit.

100 * rate({"system.network.io", "network.io.direction"="receive", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / on ("@resource.aws.elasticache.node.id") {"system.network.bandwidth.limit", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

Inbound traffic is measured against the whole limit, not half of it.

sum by ("network.allowance_exceeded.reason") ( rate({"system.network.allowance_exceeded", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

An exceedance means packets were queued or dropped.

rate({"system.network.io", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

To compare against the limit, use Network utilization vs the instance baseline.

100 * histogram_quantile(0.99, {"system.network.io.rate", "@resource.aws.elasticache.replication_group.id"="my-cluster"}) / on ("@resource.aws.elasticache.node.id") group_left {"system.network.bandwidth.limit", "@resource.aws.elasticache.replication_group.id"="my-cluster"}

The p99 of the distribution catches bursts that an averaged value smooths away.

histogram_quantile(0.99, {"system.network.io.rate", "network.io.direction"="receive", "@resource.aws.elasticache.replication_group.id"="my-cluster"})

Reveals one-second bursts that an average hides.

histogram_quantile(0.99, {"system.network.packet.rate", "network.io.direction"="receive", "@resource.aws.elasticache.replication_group.id"="my-cluster"})
rate({"system.network.packet.count", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Packet rate limits are separate from bandwidth limits. Exceeding one is reported as a pps allowance exceedance.

Requires detailed monitoring: valkey.network.output.replication

sum by ("@resource.aws.elasticache.node.id") ( rate({"valkey.network.output.replication", "valkey.role"="primary", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

This counts the bytes sent to every replica, so it scales with the number of replicas. For write throughput, see Write throughput on each primary.

Requires detailed monitoring: valkey.network.output, valkey.network.output.replication

rate({"valkey.network.output.replication", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / rate({"valkey.network.output", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Requires detailed monitoring: valkey.network.input

rate({"valkey.network.input", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[1m])

Compare against host-level network metrics to separate engine traffic from total traffic.

Errors and rejections

sum by ("valkey.error.type") ( rate({"valkey.errors", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))
topk(5, sum by ("valkey.error.type") ( rate({"valkey.errors", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])))

Requires detailed monitoring: valkey.command.rejected_calls

sum by ("valkey.command") ( rate({"valkey.command.rejected_calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Requires detailed monitoring: valkey.command.failed_calls

sum by ("valkey.command") ( rate({"valkey.command.failed_calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Requires detailed monitoring: valkey.command.failed_calls

sum by ("valkey.command.category") ( rate({"valkey.command.failed_calls", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Access control and security

sum by ("valkey.acl.denied.reason") ( rate({"valkey.acl.denied", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

A denial is either a misconfigured client or an unauthorized attempt.

sum(rate({"valkey.acl.denied", "valkey.acl.denied.reason"="auth", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Sustained failures indicate either a credential problem or a brute-force attempt.

Event loop

Requires detailed monitoring: valkey.eventloop.duration

rate({"valkey.eventloop.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

The fraction of time that the main thread is busy. The event loop is the engine's serialization point, so latency rises sharply as this approaches 1.

Requires detailed monitoring: valkey.eventloop.command.duration, valkey.eventloop.duration

rate({"valkey.eventloop.command.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / rate({"valkey.eventloop.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

A falling share points to engine overhead such as replication or defragmentation.

Requires detailed monitoring: valkey.eventloop.cycles, valkey.eventloop.duration

1e6 * rate({"valkey.eventloop.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) / rate({"valkey.eventloop.cycles", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

1e6 converts seconds to microseconds.

Requires detailed monitoring: valkey.eventloop.cycles

rate({"valkey.eventloop.cycles", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m])

Read with the average cycle time: a falling cycle rate together with rising cycle duration is the signature of a stall. An idle node still shows a floor from timer activity.

Requires detailed monitoring: valkey.eventloop.duration

rate({"valkey.eventloop.duration", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]) - on ("@resource.aws.elasticache.node.id") sum without ("cpu.mode") ( rate({"process.cpu.time", "thread.type"="main", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Event-loop duration is wall-clock time and CPU time is time on the processor, so the two normally track each other. A sustained gap means that the main thread is busy but waiting on something other than CPU, such as swap or a fork for a background save.

Fleet-wide and cross-resource patterns

topk(10, {"valkey.memory.used", "@resource.aws.elasticache.replication_group.id"=~"prod-.*"})
sum by ("@resource.aws.elasticache.replication_group.id") ( rate({"valkey.commands.processed", "@resource.aws.elasticache.replication_group.id"=~"prod-.*"}[5m]))
topk(5, sum by ("@resource.aws.elasticache.replication_group.id") ( rate({"valkey.errors", "@resource.aws.elasticache.replication_group.id"=~"prod-.*"}[5m])))
sum by ("@resource.cloud.availability_zone_id", "network.io.direction") ( rate({"system.network.io", "@resource.aws.elasticache.replication_group.id"=~"prod-.*"}[5m]))
sum by ("@resource.aws.elasticache.shard.id") ( deriv({"valkey.replication.offset", "valkey.role"="primary", "@resource.aws.elasticache.replication_group.id"="my-cluster"}[5m]))

Compare with TPS of each shard to see whether write traffic and total traffic are balanced the same way.