Quickstart: Send metrics, logs, and traces to CloudWatch with OpenTelemetry
In about 20 minutes, you send three signals from one application to CloudWatch: a metric, a log record, and a trace. You see the metric in Query Studio, the log in CloudWatch Logs, and the trace in Transaction Search.
This is the recommended path for new telemetry. You use the OpenTelemetry SDK, then the OpenTelemetry Collector, then the CloudWatch OpenTelemetry Protocol (OTLP) endpoints, and finally Query Studio, CloudWatch Logs, and Transaction Search. This path works the same way on Amazon EKS, Amazon ECS, Amazon EC2, and on-premises servers. The instrumentation that you write here doesn't change when you deploy.
When you finish, you have the following:
-
A counter metric and a histogram metric, a log record, and a trace with a span, all emitted by a sample application
-
A Collector that signs and forwards all three signals to CloudWatch
-
A PromQL query and an alarm for the metric, the log record in CloudWatch Logs, and the trace in Transaction Search
The resources that you create in this quickstart might incur charges to your AWS account. Ingesting metrics, logs, and traces can each incur charges. For more information, see OTel metrics pricing and storage.
Choosing the right path
Use this quickstart for any new custom metric. Use a different approach only if one of the following applies to you.
| If you | Use instead |
|---|---|
Run in an AWS Region where OTLP metrics ingestion isn't available. For more information about supported Regions, see Supported AWS Regions. |
|
Already have alarms and dashboards on an existing custom namespace that you must keep |
Keep publishing with |
Need a metric from existing log events without changing code |
A metric filter on the log group. For more information, see Creating metrics from log events using filters in the Amazon CloudWatch Logs User Guide. |
Want request rate, errors, and latency without writing code |
Application Signals. Add custom metrics only for business events. For more information, see Application Signals. |
Prerequisites
Before you begin this quickstart, make sure that you have the following:
-
An AWS account and access to Query Studio in the CloudWatch console
in a supported AWS Region. This quickstart uses aa-example-1. Replace it with your Region everywhere that it appears. -
AWS credentials that the Collector can find. On AWS compute, use an AWS Identity and Access Management (IAM) role. For example, use an instance profile on Amazon EC2, use IAM roles for service accounts or Amazon EKS Pod Identity on Amazon EKS, or use a task role on Amazon ECS. When you test locally, use environment variables. To get credentials for the AWS CLI, see Getting set up with the AWS CLI in the AWS Command Line Interface User Guide.
-
CloudWatch Transaction Search enabled in your Region. You can send traces to the OTLP traces endpoint only after you enable Transaction Search. For more information, see Transaction Search.
-
Docker, to run the Collector.
-
Python 3.9 or later, for the sample application. The Java, Go, .NET, and Node.js SDKs work the same way.
Important
Traces that you send to the X-Ray OTLP endpoint require Transaction Search to be enabled. This is a one-time, account-wide setting. It switches all X-Ray span ingestion into CloudWatch Logs collection mode, billed under CloudWatch pricing, with 1 percent of spans indexed for free. Metrics and logs don't require this setting.
To enable Transaction Search, run the following command.
aws xray update-trace-segment-destination --destination CloudWatchLogs --region aa-example-1
To verify, run the following command. The destination is ready when
Status is ACTIVE and Destination
is CloudWatchLogs.
aws xray get-trace-segment-destination --region aa-example-1
IAM policy for the Collector's role — The
Collector needs permissions to send metrics, logs, and traces. Attach the AWS managed
policy CloudWatchAgentServerPolicy to the Collector's role. This is the
simplest option, and it grants everything that the Collector needs for all three
signals.
Note
Sending metrics over OTLP uses the cloudwatch:PutMetricData
permission, even though you never call that API. CloudWatch metrics don't support
resource-level permissions, so policies that allow this action set
Resource to "*". The
CloudWatchAgentServerPolicy managed policy already includes this
permission, along with the permissions for logs and traces.
People who query metrics or create alarms also need the
cloudwatch:GetMetricData and cloudwatch:ListMetrics
permissions. For more information, see IAM permissions for PromQL.
Workloads outside AWS can use a bearer token instead of SigV4 for metrics and logs. The traces endpoint doesn't support bearer tokens; traces require SigV4. For more information, see Setting up bearer token authentication for Metrics.
Step 1: Instrumenting your application
Your application sends telemetry to a local Collector, not directly to CloudWatch. The SDK's OTLP exporter doesn't sign requests with SigV4, so CloudWatch rejects direct calls to the endpoints. The Collector handles signing, batching, and retries.
The following command installs the SDK and the OTLP HTTP exporter. The HTTP exporter package provides the metric, span, and log exporters that this quickstart uses.
pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-http
Create app.py. It records one counter (orders placed) and one
histogram (checkout latency), wraps each checkout in a trace span, and writes a log
record for each order. All three signals go to the local Collector.
import logging, random, time from opentelemetry import metrics, trace from opentelemetry._logs import set_logger_provider from opentelemetry.sdk.resources import Resource from opentelemetry.sdk.metrics import MeterProvider from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.sdk._logs import LoggerProvider, LoggingHandler from opentelemetry.sdk._logs.export import BatchLogRecordProcessor from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter from opentelemetry.exporter.otlp.proto.http._log_exporter import OTLPLogExporter # Static metadata goes on the resource, not on every data point resource = Resource.create({ "service.name": "checkout", "service.version": "1.4.0", "deployment.environment": "dev", }) # Metrics: local Collector; exports every 60 seconds metric_exporter = OTLPMetricExporter(endpoint="http://localhost:4318/v1/metrics") reader = PeriodicExportingMetricReader(metric_exporter, export_interval_millis=60000) metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader])) # Traces: local Collector trace.set_tracer_provider(TracerProvider(resource=resource)) trace.get_tracer_provider().add_span_processor( BatchSpanProcessor(OTLPSpanExporter(endpoint="http://localhost:4318/v1/traces")) ) # Logs: local Collector, bridged from the Python logging module logger_provider = LoggerProvider(resource=resource) logger_provider.add_log_record_processor( BatchLogRecordProcessor(OTLPLogExporter(endpoint="http://localhost:4318/v1/logs")) ) set_logger_provider(logger_provider) logging.getLogger().addHandler(LoggingHandler(logger_provider=logger_provider)) logging.getLogger().setLevel(logging.INFO) meter = metrics.get_meter("checkout") orders = meter.create_counter("orders_placed", unit="1", description="Orders placed") latency = meter.create_histogram("checkout_duration", unit="s", description="Checkout latency") tracer = trace.get_tracer("checkout") log = logging.getLogger("checkout") while True: try: payment = random.choice(["card", "wallet"]) with tracer.start_as_current_span("checkout") as span: span.set_attribute("payment_method", payment) orders.add(1, {"payment_method": payment}) duration = random.uniform(0.05, 1.2) latency.record(duration, {"payment_method": payment}) log.info("Order placed using %s in %.3fs", payment, duration) time.sleep(1) except Exception as error: print(f"Error recording telemetry: {error}")
Important
Use attributes whose values come from a small, fixed set, such as
payment_method. Never use order IDs, user IDs, or URLs that contain
IDs as attribute values.
Step 2: Running the Collector
In this quickstart, you run the Collector locally and give it credentials through environment variables. In production, the Collector runs on your compute platform and uses an IAM role instead, as shown in Next steps.
Important
Sending metrics, logs, and traces to CloudWatch incurs charges. For more information, see OTel metrics pricing and storage.
The Collector receives all three signals from your application and exports each one
to its CloudWatch OTLP endpoint. Use a distribution that includes the sigv4auth
extension, such as the OpenTelemetry Collector Contrib distribution or the AWS Distro
for OpenTelemetry (ADOT) Collector.
Note
The CloudWatch agent is an alternative. It can receive OpenTelemetry Protocol (OTLP) data and send it to CloudWatch, so you can use it instead of running a separate OpenTelemetry Collector. For more information, see Collect metrics and traces with OpenTelemetry.
The logs exporter sends to the log group and log stream that you set in the
x-aws-log-group and x-aws-log-stream headers. This
example uses a /otel/quickstart log group and a default
stream. Create them in CloudWatch Logs first, or change the values to a log group and stream
that already exist.
Create collector.yaml. In the following example, replace
aa-example-1 with your Region.
extensions: sigv4auth/metrics: service: monitoring # always "monitoring" for metrics region: aa-example-1 sigv4auth/logs: service: logs # always "logs" for logs region: aa-example-1 sigv4auth/traces: service: xray # always "xray" for traces region: aa-example-1 receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: memory_limiter: check_interval: 1s limit_percentage: 80 # Valid range: 1-100 spike_limit_percentage: 20 # Valid range: 1-100 batch: send_batch_size: 200 # Number of records to buffer before sending send_batch_max_size: 1000 # Maximum batch size; CloudWatch accepts at most 1,000 metric data points for each request timeout: 10s # Maximum time to wait before sending a batch exporters: otlphttp/metrics: # the endpoints support HTTP only, not gRPC metrics_endpoint: https://monitoring.aa-example-1.amazonaws.com/v1/metrics compression: gzip auth: authenticator: sigv4auth/metrics otlphttp/logs: # the endpoints support HTTP only, not gRPC logs_endpoint: https://logs.aa-example-1.amazonaws.com/v1/logs compression: gzip headers: x-aws-log-group: /otel/quickstart # create this log group, or use an existing one x-aws-log-stream: default # create this stream, or use an existing one auth: authenticator: sigv4auth/logs otlphttp/traces: # the endpoints support HTTP only, not gRPC traces_endpoint: https://xray.aa-example-1.amazonaws.com/v1/traces compression: gzip auth: authenticator: sigv4auth/traces service: extensions: [sigv4auth/metrics, sigv4auth/logs, sigv4auth/traces] pipelines: metrics: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlphttp/metrics] logs: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlphttp/logs] traces: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlphttp/traces]
The logs endpoint writes to an existing log group and log stream. Before you start
the Collector, create the log group and stream named in the x-aws-log-group
and x-aws-log-stream headers of your collector.yaml
logs exporter. Otherwise, CloudWatch rejects your logs with The specified log group
does not exist.
aws logs create-log-group --log-group-name /otel/quickstart --region aa-example-1 aws logs create-log-stream --log-group-name /otel/quickstart --log-stream-name default --region aa-example-1
Run the Collector locally with your credentials, and then start the application:
docker run --rm -p 4317:4317 -p 4318:4318 \ -v "$PWD/collector.yaml:/etc/otelcol-contrib/config.yaml" \ -e AWS_ACCESS_KEY_ID -e AWS_SECRET_ACCESS_KEY -e AWS_SESSION_TOKEN \ otel/opentelemetry-collector-contrib:latest python app.py
In production, pin a Collector version instead of using latest. On
AWS compute, omit the -e flags. The Collector uses the role's
credentials.
Step 3: Checking that your telemetry arrived
Data usually appears 1–2 minutes after the first export. Verify each signal in turn.
Metrics — In the CloudWatch console, open Query Studio and run each of the following queries. For more information, see Running a PromQL query in Query Studio.
| To | PromQL |
|---|---|
Confirm that the series exists |
|
Show orders per second, by payment method |
|
Show 95th percentile checkout latency |
|
The raw orders_placed query returns one series for each payment method.
Because a counter only increases, each series shows a cumulative total, as in the
following example.
orders_placed{payment_method="card"} 512 orders_placed{payment_method="wallet"} 488
The sum by (payment_method) (rate(orders_placed[5m])) query converts
those totals to a per-second rate. With the sample application, each series is near 0.5
per second, as in the following example.
{payment_method="card"} 0.51 {payment_method="wallet"} 0.49
Counters need rate() or increase(), because a raw counter
only increases. You can query gauges directly. CloudWatch stores histograms as native
histograms, so you query the metric name directly, without _bucket series
or an le label. For more information, see Querying histogram metrics.
Logs — In the CloudWatch console, open CloudWatch Logs and
open the /otel/quickstart log group (or the log group that you set in the
x-aws-log-group header). Each order writes one log record. To search
the records, use CloudWatch Logs Insights. For more information, see Analyze log data with CloudWatch Logs Insights in the Amazon CloudWatch
Logs User Guide.
In Logs Insights, select the /otel/quickstart log group and run the
following query to find your records:
fields @timestamp, severityText, body, traceId, spanId | filter @message like /checkout/ | sort @timestamp desc | limit 50
CloudWatch stores the OpenTelemetry log record as JSON, with resource attributes nested
under resource.attributes (for example,
resource.attributes.service.name), and the span's
payment_method attribute is not a field on the log record — it
appears in the span and in the log body. If you need to filter by a nested field, run a
bare fields @timestamp, @message | limit 50 query first to see the exact
field paths.
Traces — In the CloudWatch consolecheckout span. If you don't see traces, confirm that Transaction Search
is enabled in your Region.
Transaction Search stores spans in the aws/spans log group. You can
search them in the CloudWatch Transaction Search console, filtered by the checkout
service, or with a Logs Insights query on aws/spans. In Logs Insights,
select the aws/spans log group and run the following query. This query
filters on the top-level span name to avoid nested-field-path issues.
fields @timestamp, name, durationNano, attributes.payment_method, status.code, traceId | filter name = "checkout" | sort @timestamp desc | limit 20
Each span carries the trace and span IDs that also appear on your log records, so you
can pivot between a log line and its trace. CloudWatch enriches spans with Application Signals
attributes such as aws.local.service.
If you don't see data after 5 minutes, see Troubleshooting telemetry export.
Step 4: Creating an alarm and a dashboard widget
Turn the latency query into an alarm and a dashboard widget, both from Query Studio.
To create an alarm and a dashboard widget
-
Open the CloudWatch console
, and then open Query Studio. -
In Query Studio, run the 95th percentile latency query from Step 3: Checking that your telemetry arrived.
-
From the actions menu, choose Create alarm. For more information, see Creating alarms from Query Studio.
Note
If you receive an
AccessDeniederror, verify that theCloudWatchAgentServerPolicymanaged policy is attached, or that your IAM policy includes thecloudwatch:GetMetricDataandcloudwatch:ListMetricspermissions. For the required permissions, see Prerequisites. -
Set the condition to greater than
0.8seconds for 3 out of 3 periods. -
Choose an Amazon SNS topic for notifications.
-
Return to Query Studio, choose Add to dashboard, and choose or create a dashboard for the checkout service. For more information, see Adding visualizations to dashboards.
-
(Optional) Add
sum(rate(orders_placed[5m]))to the same dashboard. -
(Optional) Turn on anomaly detection for the orders metric, so that a drop in orders alerts you without a fixed threshold. For more information, see Anomaly detection using PromQL.
Troubleshooting telemetry export
Check the Collector logs first. Export errors include the HTTP status code from CloudWatch.
| Symptom | Likely cause | Fix |
|---|---|---|
403 on metrics in the Collector logs |
The role lacks |
Attach |
403 on logs in the Collector logs |
The role lacks the logs permissions, or the
|
Attach |
4xx on logs, missing header |
The logs exporter doesn't set the
|
Set both headers on the |
A log event is truncated or rejected |
A single request to the logs endpoint is larger than 1 MB. |
Reduce the batch size or the size of each log record so that each request stays under 1 MB. |
Traces rejected |
Transaction Search isn't enabled in the Region. |
Enable CloudWatch Transaction Search, and then resend the traces. |
403 on traces in the Collector logs |
A bearer token was used for traces. The traces endpoint doesn't support bearer tokens. |
Use SigV4 for traces. Configure the |
Connection errors to CloudWatch |
The exporter is gRPC ( |
Use |
The application logs |
The Collector isn't running, or port 4318 isn't published. |
Start the Collector, and check |
400, too many data points |
A batch has more than 1,000 data points. |
Set |
400 on some series |
A data point has more than 150 labels, or a label value is longer than 1,024 characters. |
Drop or shorten attributes in a |
429 |
More than 500 requests per second for the account, or too many new series. |
Send larger batches, and reduce the number of attribute values. |
Accepted, but not visible |
The data point timestamp is more than 14 days in the past or more than 2 hours in the future. |
Check the host clock. |
For all limits, see Endpoint limits and restrictions.
Clean up resources
To avoid ongoing charges, delete the resources that you created in this quickstart when you no longer need them:
-
Delete the alarm that you created.
-
Remove the dashboard widget that you added.
-
Delete the
/otel/quickstartlog group and itsdefaultlog stream (or the log group and stream that you used) in CloudWatch Logs. -
Stop the Collector container and the sample application.
-
Detach
CloudWatchAgentServerPolicyfrom the Collector's role if you attached it only for this quickstart.
Keeping costs predictable
CloudWatch bills OTLP metrics per GB ingested, with 15 months of storage included. Logs and traces are also billed by the volume that you ingest, and Transaction Search has its own charges. Volume grows with the number of data points per series, the number of series, the number of log records, and the number of traces. For more information, see OTel metrics pricing and storage.
To keep costs predictable, keep the following practices in mind:
-
Export every 60 seconds. Halving the interval roughly doubles the volume.
-
Count your series before you ship. The number of series for a metric is the product of the number of distinct values of each attribute. For example, 2 payment methods, 3 Regions, and 5 status codes produce 30 series. Adding a user ID makes the number unbounded.
-
Put static metadata on the resource. As shown in Step 1: Instrumenting your application, service name, version, and environment belong on the resource, not on each data point.
-
Use histograms where you need distributions. A histogram data point carries more data than a counter or gauge data point. Use histograms for latency and sizes, not for everything.
-
Drop what you don't query. With a
filterprocessor in the Collector, you can remove unused metrics, log records, spans, or attributes before you're billed for them.
Next steps
Your instrumentation stays the same in production. Only where the Collector runs changes.
The following table shows where the Collector runs and which credentials to use in each environment.
| Environment | Where the Collector runs | Credentials |
|---|---|---|
Amazon EKS |
A DaemonSet or Deployment. Applications send to its Service. |
IAM roles for service accounts or Amazon EKS Pod Identity |
Amazon ECS |
A sidecar container in the task |
Task role |
Amazon EC2 |
A service on each instance |
Instance profile |
On-premises |
A service on each host, or a central gateway |
IAM Roles Anywhere, or a bearer token for metrics and logs |
Before you go to production, also plan for the following:
-
Configure the sending queue and retry settings on each exporter.
-
Set up Collector health checks.
-
Set up cross-account observability if telemetry from many accounts goes to one monitoring account. For more information, see CloudWatch cross-account observability.
For more ways to send telemetry, see Collect and send telemetry.