Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/en/changes/changes.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
## 11.1.0

#### Project
* Update the e2e mock sender's OTLP proto to v1.11.1. A current OpenTelemetry Collector adds fields the old copy lacked, such as a metric's `metadata`, and the sender's strict JSON parser refused a recording that carries them. The Windows e2e now also replays such a recording, of windows_exporter `0.31.8` through OpenTelemetry Collector `0.158.0` taken on a GitHub-hosted Windows runner, and asserts its memory metrics.
* Replace the GenAI e2e cases' Spring AI application with an in-repo `e2e-spring-ai-service` module pinned to the released Spring AI 2.0.1, and move the mock LLM endpoints out of `e2e-service-provider` into a dedicated `e2e-mock-llm-server` module that all three GenAI cases share. The Spring AI application was previously built at test time by cloning `spring-projects/spring-ai-examples` and running Maven inside the image build, which resolved `spring-ai:2.0.0-SNAPSHOT`: a line that has since been abandoned, whose surviving builds all predate the 2.0.0 GA by a day and whose older builds have been pruned, so the fixture was pinned to nothing released and could no longer be rebuilt from the same bytes. Both modules are built by the existing e2e reactor and published to `ghcr.io/apache/skywalking` alongside the other e2e service images; `e2e-spring-ai-service` needs Java 17 and Spring Boot 4, so it is gated behind a new `jdk-17` profile in the e2e reactor and the jobs that build it now pin JDK 17 explicitly; `make -C test docker` refuses up front on an older JDK rather than packaging a missing or stale jar, since `clean` never reaches a module the profile excluded and compose would turn the absent jar into a directory.

#### OAP Server
Expand Down Expand Up @@ -28,13 +29,16 @@
* Bring the TraceQL module closer to the Tempo API. Every datasource lists only the spans that matched a search in each span set, capped at `spss`, fills `serviceStats`, filters `span.http.status_code`, honours `scope` and the per-scope `limit` on the tag endpoints, whose `intrinsic` scope lists what each datasource can filter, accepts Tempo's unscoped `.key` tag names on the tag-value endpoints, and answers a tag-name or tag-value lookup filtered by `q` from a sample of the newest 50 matching traces, so Grafana's dependent dropdowns narrow down; the `status` and `kind` dropdowns always list their full enum. The `/otlp` datasource also accepts Tempo's `span:name`, `span:kind`, `span:status` and `span:duration` spellings. `/api/metrics/query_range` and `/api/metrics/query` answer `501 Not Implemented` with an error body and accept Tempo's `q` parameter, instead of `200` with a string Grafana cannot parse.
* Refuse unsupported TraceQL with `400` instead of dropping the predicate and answering with unfiltered traces: syntax errors (`||` inside a spanset, regex, an attribute without a scope or leading dot), `!=` and other unsupported operators, negation, attribute-existence checks, multiple spansets, unknown intrinsics and values, and conditions a datasource cannot filter (`kind` on Zipkin and SkyWalking, `resource.instance` on Zipkin, `resource.remote.service` and `status = unset` on SkyWalking, `resource.instance` or `name` without a service on SkyWalking). The deprecated `tags` search parameter is parsed as logfmt on every datasource and goes through the same mapping and refusals. `kind = server` and `status = error` accept the bare keyword as Tempo does. Every search-result span carries a `status` attribute (`error`, `ok`, `unset`) next to `service.name` and `span.kind`, so a trace list can show failures. A trace the storage matched but none of whose spans satisfies every condition is left out instead of listed with every span, and the `/otlp` datasource compares names the way the receiver indexed them and keeps the `resource.` and `span.` scopes apart in its tag index, so `span.env` no longer matches a resource attribute. `duration >` and `<` are strict at microsecond precision. (#14093)
* Support the OpenTelemetry Collector `hostmetrics` receiver as an alternative source for Linux and Windows host monitoring. `vm.yaml` and `windows.yaml` map it to the same `meter_vm_*` / `meter_win_*` metrics as node-exporter and windows_exporter, whose metrics keep their meaning, and add metrics only hostmetrics provides: CPU core count and normalized CPU usage, plus CPU load, file-system usage, system handle count and pagefile usage on Windows. node-exporter metrics without a matching hostmetrics source (`tcp_alloc`, `sockets_used`, `udp_inuse`, `filefd_allocated`) stay node-exporter only. The new `process-hostmetrics-linux` and `process-hostmetrics-windows` rules, enabled by default, report each process name as an instance of its host (`mp_process_linux_*` / `mp_process_windows_*`: process count, threads, CPU, resident memory, open handles and oldest process uptime). Reference Collector configurations for both systems are under `docs/en/setup/backend/`; they set the `job_name` and normalized `process_name` labels the rules route and group on, and sum same-named processes before export.
* Windows monitoring through windows_exporter now requires windows_exporter `0.29.0` or later. windows_exporter `0.31.0` removed the `cs` collector and the deprecated `os` memory metrics (`windows_cs_physical_memory_bytes`, `windows_os_physical_memory_free_bytes`, `windows_os_virtual_memory_bytes`, `windows_os_virtual_memory_free_bytes`) that `windows.yaml` read, so on `0.31.0`+ `meter_win_memory_total` / `_available` / `_used` and the `meter_win_memory_virtual_memory_*` commit metrics were empty. The rules now read their replacements from the `memory` collector (`windows_memory_physical_total_bytes`, `windows_memory_physical_free_bytes`, `windows_memory_commit_limit`, `windows_memory_committed_bytes`), all of which windows_exporter provides since `0.29.0`; on an older windows_exporter these memory metrics stay empty until it is upgraded. Also fix `meter_win_cpu_total_percentage` counting overlapping Windows CPU modes twice: windows_exporter's `privileged` already includes `interrupt` and `dpc` time, and OpenTelemetry hostmetrics' `system` already includes `interrupt` time, so busy CPU is now `user` + `privileged` or `user` + `system`.
* Fix the OpenTelemetry receiver creating no Linux or Windows host service from an OpenTelemetry Collector `0.127.0` or later. The Collector's Prometheus receiver sends a scraped target's host only as `server.address` from that release, no longer as `net.host.name`, and `vm.yaml` and `windows.yaml` name a host by `node_identifier_host_name`, which the receiver took only from `net.host.name` or `host.name`. It now falls back to `server.address` when neither is present.
* AI agent conversations adopt AI Sessionizer `130601c`. Calls to MCP servers: the `execution` kind the Sessionizer's Claude Code plugin writes, `streams/<stream>/execution-<stamp>-<seq>.sd`, is stored like any other file; the document lists every `execution/1` record under `tool_executions`, joined to its step by tool-use id and kept once by its own id, a tool step names its records under `executions`, and a call to an MCP server carries `mcp_server` and `mcp_tool` in its `attrs`. The new `otel-rules/ai-agent/mcp_endpoint.yaml` gives one endpoint per MCP server and tool, `<server>/<tool>`, under the agent's service, with `meter_ai_agent_mcp_calls`, `meter_ai_agent_mcp_calls_by_outcome` and `meter_ai_agent_mcp_duration` from the Sessionizer's `agent.mcp.calls` and `agent.mcp.duration`. The token rules read `agent.token.usage`, the Sessionizer's name for the metric. The document follows the Sessionizer's: an `llm.call` takes its provider bodies from the round's `provider_bodies` attribute and the session its count from `provider_bodies_landed`, and a round whose bodies do not read or point past its range is refused; talks are in the order they began across streams; only a `child` stream's talk is a child's, not an `auxiliary` one's; a stream's `opened_by`, the relations and a step's edges are in the order they happened, by record position inside one stream or workflow run and by time across them, never by id; change and execution records of one instant are in the order they were read; a record is read by the fields its format lists, and one of another shape is skipped; a round with a frame field of another type, or a negative sequence, row or round number, does not read; a record time is RFC 3339 as Session Data defines it, to the second, and anything else is no time; record times compare as instants.
* Fix the Elasticsearch `BulkProcessor` treating a bulk write as fully successful whenever the overall HTTP status was 200, even when Elasticsearch rejected individual items in the same response, for example a 429 `cluster_block_exception` when an index is switched to `read_only_allow_delete` by the flood-stage disk watermark. The response body's `errors` flag and per-item `status`/`error` are now decoded and checked on every 200 response; a rejected request's future now completes exceptionally instead of being treated as a success, while the rest of the batch still completes normally. One ERROR line per `_bulk` request summarizes the rejections grouped by status code and error type, never a document id or the raw ES error reason. A response whose `items` array is shorter than the request (a malformed or truncated response) also fails those requests instead of silently treating them as succeeded. This also fixes a related batching bug where, once a flush was split into multiple `_bulk` HTTP requests by `batchOfBytes`, the completion of one chunk's request completed or failed every request in the whole flush instead of just its own chunk, and a bug where requests whose bulk failed to even build (e.g. an encoding error) were left with a future that never completed, which could block a persistence round forever.

#### UI
* Add a Virtual GenAI evaluation-record page and evaluation-score chart in Horizon UI, so operators can inspect evaluation result, level, reason, judge model, timestamp, trace linkage, and the `gen_ai_model_evaluation_score_ppm` trend for evaluated records.

#### Documentation
* The Windows monitoring doc's example OpenTelemetry Collector configuration uses the `debug` exporter. The `logging` exporter it used was removed from the Collector, so the example no longer loaded.
* Document the BanyanDB trace tail sampling metrics in the BanyanDB self-observability dashboard catalog, and point the "Operating it" section of the trace tail sampling guide at them — the OAP-collected metrics show what the sampler plugins *proposed* next to what storage *committed*, which the data node's raw metrics endpoint alone does not.
* Rename Envoy AI Gateway to Agent Router in the monitoring docs, following the project's move to the Agentic AI Foundation. Only the prose changes: the `ENVOY_AI_GATEWAY` layer, the `job_name=envoy-ai-gateway` routing tag, the `envoy-ai-gateway` rule files and the `meter_envoy_ai_gw_` metric names are kept, as Agent Router itself kept its deployed names and the telemetry it emits.

Expand Down
2 changes: 1 addition & 1 deletion docs/en/setup/backend/backend-vm-monitoring.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ Likewise, `meter_vm_tcp_alloc`, `meter_vm_sockets_used`, and `meter_vm_udp_inuse
| TCP Established / Close-Wait | count | `meter_vm_tcp_curr_estab` | TCP connections in ESTABLISHED or CLOSE-WAIT state | Yes | Yes | Yes |
| TCP Time Wait | count | `meter_vm_tcp_tw` | TCP connections in TIME-WAIT state | Yes | Yes | Yes |
| TCP Allocated | count | `meter_vm_tcp_alloc` | Allocated TCP sockets | Yes | No | Yes |
| Sockets Used | count | `meter_vm_sockets_used` | Kernel sockets currently in use | Yes | No | Yes |
| Sockets Used | count | `meter_vm_sockets_used` | Kernel sockets currently in use | Yes | No | — |
| UDP In Use | count | `meter_vm_udp_inuse` | UDP sockets currently in use | Yes | No | Yes |
| Filefd Allocated | count | `meter_vm_filefd_allocated` | Host-level allocated file descriptors from Linux `/proc/sys/fs/file-nr` | Yes | No | — |

Expand Down
10 changes: 7 additions & 3 deletions docs/en/setup/backend/backend-win-monitoring.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,14 @@ Windows entity as a `Service` in OAP and on the `Layer: OS_WINDOWS`.
3. The SkyWalking OAP Server parses the expression with [MAL](../../concepts-and-designs/mal.md) to filter/calculate/aggregate and store the results.
## Setup
**For OpenTelemetry receiver:**
1. Setup [Prometheus windows_exporter](https://github.com/prometheus-community/windows_exporter).
1. Setup [Prometheus windows_exporter](https://github.com/prometheus-community/windows_exporter) `0.29.0` or later.
- Keep its default `cpu`, `memory`, `logical_disk` and `net` collectors enabled.
- Memory metrics are read from the `memory` collector. windows_exporter `0.31.0` removed the `cs` collector and the `os` memory metrics that older SkyWalking releases read.
2. Setup [OpenTelemetry Collector ](https://opentelemetry.io/docs/collector/). This is an example for OpenTelemetry Collector configuration [otel-collector-config.yaml](../../../../test/e2e-v2/cases/win/prometheus-windows_exporter/otel-collector-config.yaml).
3. Config SkyWalking [OpenTelemetry receiver](opentelemetry-receiver.md).

Each Windows host is a service named after its host: the `host.name` or `net.host.name` resource attribute when the Collector sends one, otherwise `server.address`, where the Collector's Prometheus receiver puts the scraped target's host from Collector `0.127.0`.

### Native OpenTelemetry hostmetrics (expanded alternative)
SkyWalking can also receive Windows host and process metrics directly from OpenTelemetry Collector Contrib, without Prometheus windows_exporter.

Expand All @@ -29,8 +33,8 @@ The `OTel hostmetrics` column below refers to the complete OpenTelemetry Collect

| Monitoring Panel | Unit | Metric Name | Description | windows_exporter | OTel hostmetrics |
|---|---|---|---|:---:|:---:|
| CPU Usage | % | `meter_win_cpu_total_percentage` | Total CPU usage across all logical CPUs | Yes | Yes |
| CPU Average Used | % | `meter_win_cpu_average_used` | CPU usage by mode/state | Yes | Yes |
| CPU Usage | % | `meter_win_cpu_total_percentage` | Busy CPU across all logical CPUs: `user` + `privileged` (windows_exporter) or `user` + `system` (OTel) | Yes | Yes |
| CPU Average Used | % | `meter_win_cpu_average_used` | CPU usage by mode/state, summed across logical CPUs. Windows modes overlap: `privileged` (windows_exporter) includes `interrupt` and `dpc` time, and `system` (OTel) includes `interrupt` time | Yes | Yes |
| CPU Cores | count | `meter_win_cpu_cores_num` | Number of logical CPUs | — | Yes |
| Normalized CPU Usage | % | `meter_win_cpu_norm_percentage` | CPU usage normalized by the number of logical CPUs | — | Yes |
| CPU Load | | `meter_win_cpu_load1`<br />`meter_win_cpu_load5`<br />`meter_win_cpu_load15` | CPU load metrics exposed by the Collector hostmetrics load scraper | — | Yes |
Expand Down
7 changes: 4 additions & 3 deletions docs/en/setup/backend/opentelemetry-receiver.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,8 @@ receiver-otel:
```

The receiver adds label with key `node_identifier_host_name` to the collected data samples,
and its value is from `net.host.name` (or `host.name` for some OTLP versions) resource attributes defined in OpenTelemetry proto,
for identification of the metric data.
and its value is from the `net.host.name` or `host.name` resource attribute, or, when neither is present,
`server.address`, for identification of the metric data.

**Label name conversion:** Dots (`.`) in attribute key names are converted to underscores (`_`) for both
resource attributes and data point (metric-level) attributes. For example, `gen_ai.token.type` becomes
Expand All @@ -40,11 +40,12 @@ in the resource attributes, the fallback is skipped.
| `service.name` | `job_name` | The [OTel Collector Prometheus Receiver](https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/receiver/prometheusreceiver/README.md) automatically converts the Prometheus `job` label to `service.name`. This fallback ensures it is available as `job_name` for MAL rule filtering. |
| `net.host.name` | `node_identifier_host_name` | Legacy: used by VM/Windows MAL rules |
| `host.name` | `node_identifier_host_name` | Legacy: used by VM/Windows MAL rules |
| `server.address` | `node_identifier_host_name` | Only when neither `net.host.name` nor `host.name` is present. The OTel Collector Prometheus Receiver sends a scraped target's host only as `server.address` from Collector `0.127.0`. |

When `job_name` is set explicitly in `OTEL_RESOURCE_ATTRIBUTES` (e.g., `job_name=envoy-ai-gateway` for [Agent Router](backend-envoy-ai-gateway-monitoring.md)),
it takes precedence and the `service.name` fallback is skipped.

**Note:** The `net.host.name` and `host.name` mappings are legacy. New integrations should use
**Note:** The `net.host.name`, `host.name` and `server.address` mappings are legacy. New integrations should use
the natural dot-to-underscore conversion (e.g., `host.name` → `host_name` in MAL rules).

**Points of one request are analysed a minute at a time**, oldest minute first. A MAL rule folds every sample of
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -62,11 +62,11 @@ input:
value: 3000.0
process_memory_utilization:
- labels: {node_identifier_host_name: test-host, process_name: java}
value: 0.01
value: 1.0
- labels: {node_identifier_host_name: test-host, process_name: java}
value: 0.02
value: 2.0
- labels: {node_identifier_host_name: test-host, process_name: java}
value: 0.03
value: 3.0
process_open_handles:
- labels: {node_identifier_host_name: test-host, process_name: java}
value: 5.0
Expand Down Expand Up @@ -110,7 +110,7 @@ expected:
mp_process_linux_memory_resident_percent:
samples:
- labels: {node_identifier_host_name: test-host, process_name: java}
value: 6.0
value: 600.0
mp_process_linux_open_handles:
samples:
- labels: {node_identifier_host_name: test-host, process_name: java}
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -62,11 +62,11 @@ input:
value: 3000.0
process_memory_utilization:
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
value: 0.01
value: 1.0
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
value: 0.02
value: 2.0
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
value: 0.03
value: 3.0
process_open_handles:
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
value: 5.0
Expand Down Expand Up @@ -110,7 +110,7 @@ expected:
mp_process_windows_memory_resident_percent:
samples:
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
value: 6.0
value: 600.0
mp_process_windows_open_handles:
samples:
- labels: {node_identifier_host_name: test-host, process_name: svchost.exe}
Expand Down
Loading
Loading