Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions docs/en/changes/changes.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
## 11.1.0

#### Project
* Update the e2e mock sender's OTLP proto to v1.11.1, so it replays recordings from a current OpenTelemetry Collector, which adds fields such as a metric's `metadata` that the old copy's strict JSON parser refused.
* Replace the GenAI e2e cases' Spring AI application with an in-repo `e2e-spring-ai-service` module pinned to the released Spring AI 2.0.1, and move the mock LLM endpoints out of `e2e-service-provider` into a dedicated `e2e-mock-llm-server` module that all three GenAI cases share. The Spring AI application was previously built at test time by cloning `spring-projects/spring-ai-examples` and running Maven inside the image build, which resolved `spring-ai:2.0.0-SNAPSHOT`: a line that has since been abandoned, whose surviving builds all predate the 2.0.0 GA by a day and whose older builds have been pruned, so the fixture was pinned to nothing released and could no longer be rebuilt from the same bytes. Both modules are built by the existing e2e reactor and published to `ghcr.io/apache/skywalking` alongside the other e2e service images; `e2e-spring-ai-service` needs Java 17 and Spring Boot 4, so it is gated behind a new `jdk-17` profile in the e2e reactor and the jobs that build it now pin JDK 17 explicitly; `make -C test docker` refuses up front on an older JDK rather than packaging a missing or stale jar, since `clean` never reaches a module the profile excluded and compose would turn the absent jar into a directory.

#### OAP Server
Expand Down Expand Up @@ -29,6 +30,9 @@
* Refuse unsupported TraceQL with `400` instead of dropping the predicate and answering with unfiltered traces: syntax errors (`||` inside a spanset, regex, an attribute without a scope or leading dot), `!=` and other unsupported operators, negation, attribute-existence checks, multiple spansets, unknown intrinsics and values, and conditions a datasource cannot filter (`kind` on Zipkin and SkyWalking, `resource.instance` on Zipkin, `resource.remote.service` and `status = unset` on SkyWalking, `resource.instance` or `name` without a service on SkyWalking). The deprecated `tags` search parameter is parsed as logfmt on every datasource and goes through the same mapping and refusals. `kind = server` and `status = error` accept the bare keyword as Tempo does. Every search-result span carries a `status` attribute (`error`, `ok`, `unset`) next to `service.name` and `span.kind`, so a trace list can show failures. A trace the storage matched but none of whose spans satisfies every condition is left out instead of listed with every span, and the `/otlp` datasource compares names the way the receiver indexed them and keeps the `resource.` and `span.` scopes apart in its tag index, so `span.env` no longer matches a resource attribute. `duration >` and `<` are strict at microsecond precision. (#14093)
* Support the OpenTelemetry Collector `hostmetrics` receiver as an alternative source for Linux and Windows host monitoring. `vm.yaml` and `windows.yaml` map it to the same `meter_vm_*` / `meter_win_*` metrics as node-exporter and windows_exporter, whose metrics keep their meaning, and add metrics only hostmetrics provides: CPU core count and normalized CPU usage, plus CPU load, file-system usage, system handle count and pagefile usage on Windows. node-exporter metrics without a matching hostmetrics source (`tcp_alloc`, `sockets_used`, `udp_inuse`, `filefd_allocated`) stay node-exporter only. The new `process-hostmetrics-linux` and `process-hostmetrics-windows` rules, enabled by default, report each process name as an instance of its host (`mp_process_linux_*` / `mp_process_windows_*`: process count, threads, CPU, resident memory, open handles and oldest process uptime). Reference Collector configurations for both systems are under `docs/en/setup/backend/`; they set the `job_name` and normalized `process_name` labels the rules route and group on, and sum same-named processes before export.
* AI agent conversations adopt AI Sessionizer `130601c`. Calls to MCP servers: the `execution` kind the Sessionizer's Claude Code plugin writes, `streams/<stream>/execution-<stamp>-<seq>.sd`, is stored like any other file; the document lists every `execution/1` record under `tool_executions`, joined to its step by tool-use id and kept once by its own id, a tool step names its records under `executions`, and a call to an MCP server carries `mcp_server` and `mcp_tool` in its `attrs`. The new `otel-rules/ai-agent/mcp_endpoint.yaml` gives one endpoint per MCP server and tool, `<server>/<tool>`, under the agent's service, with `meter_ai_agent_mcp_calls`, `meter_ai_agent_mcp_calls_by_outcome` and `meter_ai_agent_mcp_duration` from the Sessionizer's `agent.mcp.calls` and `agent.mcp.duration`. The token rules read `agent.token.usage`, the Sessionizer's name for the metric. The document follows the Sessionizer's: an `llm.call` takes its provider bodies from the round's `provider_bodies` attribute and the session its count from `provider_bodies_landed`, and a round whose bodies do not read or point past its range is refused; talks are in the order they began across streams; only a `child` stream's talk is a child's, not an `auxiliary` one's; a stream's `opened_by`, the relations and a step's edges are in the order they happened, by record position inside one stream or workflow run and by time across them, never by id; change and execution records of one instant are in the order they were read; a record is read by the fields its format lists, and one of another shape is skipped; a round with a frame field of another type, or a negative sequence, row or round number, does not read; a record time is RFC 3339 as Session Data defines it, to the second, and anything else is no time; record times compare as instants.
* Fix Windows monitoring with windows_exporter v0.31.0 and later, which removed the `cs` collector and the `os` collector's memory metrics. `windows.yaml` read `windows_cs_physical_memory_bytes`, `windows_os_physical_memory_free_bytes`, `windows_os_virtual_memory_bytes` and `windows_os_virtual_memory_free_bytes`, so every `meter_win_memory_*` metric of the windows_exporter source was empty. It now reads the `memory` collector's `windows_memory_physical_total_bytes`, `windows_memory_physical_free_bytes`, `windows_memory_commit_limit` and `windows_memory_committed_bytes`, which carry the same values and are in windows_exporter's default collectors since v0.29.0, now the minimum. The Windows e2e replays a recording of windows_exporter v0.31.8 through OpenTelemetry Collector 0.158.0 in place of mock data in the old names.
* Fix `meter_win_cpu_total_percentage` counting interrupt and DPC time twice for windows_exporter. Windows' `% Privileged Time`, windows_exporter's `privileged` mode, already includes `% Interrupt Time` and `% DPC Time`, so busy time is `user` plus `privileged`. On a GitHub-hosted Windows runner the five modes summed to 1.005-1.009 seconds per second per core, and `idle`, `user` and `privileged` to exactly 1.
* Fix OpenTelemetry Collector v0.127.0 and later losing the host of Prometheus-scraped Linux and Windows metrics. The Collector's Prometheus receiver sends a target's host as `server.address` and no longer `net.host.name`, so `vm.yaml` and `windows.yaml`, which name a host by `node_identifier_host_name`, received no host and created no service. The OpenTelemetry receiver now takes `node_identifier_host_name` from `server.address` when the resource carries neither `net.host.name` nor `host.name`.
* Fix the Elasticsearch `BulkProcessor` treating a bulk write as fully successful whenever the overall HTTP status was 200, even when Elasticsearch rejected individual items in the same response, for example a 429 `cluster_block_exception` when an index is switched to `read_only_allow_delete` by the flood-stage disk watermark. The response body's `errors` flag and per-item `status`/`error` are now decoded and checked on every 200 response; a rejected request's future now completes exceptionally instead of being treated as a success, while the rest of the batch still completes normally. One ERROR line per `_bulk` request summarizes the rejections grouped by status code and error type, never a document id or the raw ES error reason. A response whose `items` array is shorter than the request (a malformed or truncated response) also fails those requests instead of silently treating them as succeeded. This also fixes a related batching bug where, once a flush was split into multiple `_bulk` HTTP requests by `batchOfBytes`, the completion of one chunk's request completed or failed every request in the whole flush instead of just its own chunk, and a bug where requests whose bulk failed to even build (e.g. an encoding error) were left with a future that never completed, which could block a persistence round forever.

#### UI
Expand Down
4 changes: 3 additions & 1 deletion docs/en/setup/backend/backend-win-monitoring.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,12 @@ Windows entity as a `Service` in OAP and on the `Layer: OS_WINDOWS`.
3. The SkyWalking OAP Server parses the expression with [MAL](../../concepts-and-designs/mal.md) to filter/calculate/aggregate and store the results.
## Setup
**For OpenTelemetry receiver:**
1. Setup [Prometheus windows_exporter](https://github.com/prometheus-community/windows_exporter).
1. Setup [Prometheus windows_exporter](https://github.com/prometheus-community/windows_exporter) v0.29.0 or later, with its default collectors. The rules read its `cpu`, `memory`, `logical_disk` and `net` collectors.
2. Setup [OpenTelemetry Collector ](https://opentelemetry.io/docs/collector/). This is an example for OpenTelemetry Collector configuration [otel-collector-config.yaml](../../../../test/e2e-v2/cases/win/prometheus-windows_exporter/otel-collector-config.yaml).
3. Config SkyWalking [OpenTelemetry receiver](opentelemetry-receiver.md).

Each Windows host is a service named after its host: the resource attribute `host.name` or `net.host.name` when the Collector sends one, otherwise `server.address`, which is where the Collector's Prometheus receiver puts the scraped target's host from v0.127.0.

### Native OpenTelemetry hostmetrics (expanded alternative)
SkyWalking can also receive Windows host and process metrics directly from OpenTelemetry Collector Contrib, without Prometheus windows_exporter.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,11 +15,29 @@

script: oap-server/server-starter/src/main/resources/otel-rules/windows.yaml
input:
# windows_exporter's privileged time already includes its interrupt and dpc
# time, so only user and privileged count as busy.
windows_cpu_time_total:
- labels:
node_identifier_host_name: test-host
mode: user
value: 100.0
- labels:
node_identifier_host_name: test-host
mode: privileged
value: 60.0
- labels:
node_identifier_host_name: test-host
mode: interrupt
value: 8.0
- labels:
node_identifier_host_name: test-host
mode: dpc
value: 4.0
- labels:
node_identifier_host_name: test-host
mode: idle
value: 200.0
# Native OpenTelemetry hostmetrics contract after Collector normalization.
# Values are fixed for deterministic MAL assertions.
system_cpu_logical_count:
Expand Down Expand Up @@ -81,18 +99,23 @@ input:
device: 'C:\\pagefile.sys'
state: free
value: 70.0
windows_cs_physical_memory_bytes:
# windows_exporter's memory collector, v0.29.0 or later.
windows_memory_physical_total_bytes:
- labels:
node_identifier_host_name: test-host
value: 100.0
windows_os_physical_memory_free_bytes:
windows_memory_physical_free_bytes:
- labels:
value: 100.0
windows_os_virtual_memory_free_bytes:
node_identifier_host_name: test-host
value: 30.0
windows_memory_commit_limit:
- labels:
node_identifier_host_name: test-host
value: 100.0
windows_os_virtual_memory_bytes:
windows_memory_committed_bytes:
- labels:
value: 100.0
node_identifier_host_name: test-host
value: 40.0
windows_logical_disk_read_bytes_total:
- labels:
node_identifier_host_name: test-host
Expand All @@ -118,59 +141,79 @@ expected:
samples:
- labels:
node_identifier_host_name: test-host
value: 3500.0
value: 5000.0
meter_win_cpu_average_used:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
node_identifier_host_name: test-host
mode: idle
value: 6500.0
- labels:
node_identifier_host_name: test-host
mode: interrupt
value: 400.0
- labels:
node_identifier_host_name: test-host
mode: user
value: 3000.0
meter_win_memory_total:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
node_identifier_host_name: test-host
value: 100.0
meter_win_memory_available:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
value: 100.0
node_identifier_host_name: test-host
value: 30.0
meter_win_memory_used:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
value: 0.0
node_identifier_host_name: test-host
value: 70.0
meter_win_memory_virtual_memory_free:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
value: 100.0
node_identifier_host_name: test-host
value: 60.0
meter_win_memory_virtual_memory_total:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
node_identifier_host_name: test-host
value: 100.0
meter_win_memory_virtual_memory_percentage:
entities:
- scope: SERVICE
service: test-host
layer: OS_WINDOWS
samples:
- labels:
value: -0.0
node_identifier_host_name: test-host
value: 40.0
meter_win_disk_read:
entities:
- scope: SERVICE
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,14 @@ public class OpenTelemetryMetricRequestProcessor implements Service, MalConverte
// in resource attributes (e.g., Envoy AI Gateway), it takes precedence via putIfAbsent.
.put("service.name", "job_name")
.build();

/**
* Where a Prometheus target's host arrives from an OTel Collector v0.127.0 or later: its Prometheus receiver
* sends {@code server.address} and no longer {@code net.host.name}. Read only when no legacy host attribute
* gave {@code node_identifier_host_name}, so an explicit {@code host.name} still wins.
*/
private static final String SERVER_ADDRESS = "server.address";
private static final String NODE_IDENTIFIER_HOST_NAME = "node_identifier_host_name";
/**
* Active MAL converters, keyed by {@code "<catalog>:<rule-name>"} so boot-time entries and
* runtime-rule entries share one namespace. A runtime {@code /addOrUpdate} for a rule that
Expand Down Expand Up @@ -145,20 +153,7 @@ public void processMetricsRequest(final ExportMetricsServiceRequest requests) {
log.debug("Resource attributes: {}", request.getResource().getAttributesList());
}

// First pass: collect all resource attributes with dots replaced by underscores
final Map<String, String> nodeLabels = new HashMap<>();
for (final var it : request.getResource().getAttributesList()) {
final String key = it.getKey().replace('.', '_');
final String value = anyValueToString(it.getValue());
nodeLabels.putIfAbsent(key, value);
}
// Second pass: apply fallback mappings — only if the target key is absent
for (final var it : request.getResource().getAttributesList()) {
final String targetKey = FALLBACK_LABEL_MAPPINGS.get(it.getKey());
if (targetKey != null) {
nodeLabels.putIfAbsent(targetKey, anyValueToString(it.getValue()));
}
}
final Map<String, String> nodeLabels = nodeLabels(request.getResource().getAttributesList());

// A request is analysed a minute at a time, oldest minute first. A MAL rule folds every sample of
// an entity into one value stamped with the first sample's time, which is right for a scrape, whose
Expand Down Expand Up @@ -255,6 +250,35 @@ public void start() throws ModuleStartException {
}
}

/**
* The labels every sample of a resource carries, from its resource attributes.
*/
static Map<String, String> nodeLabels(final List<KeyValue> attributes) {
// First pass: collect all resource attributes with dots replaced by underscores
final Map<String, String> nodeLabels = new HashMap<>();
for (final var it : attributes) {
final String key = it.getKey().replace('.', '_');
final String value = anyValueToString(it.getValue());
nodeLabels.putIfAbsent(key, value);
}
// Second pass: apply fallback mappings — only if the target key is absent
for (final var it : attributes) {
final String targetKey = FALLBACK_LABEL_MAPPINGS.get(it.getKey());
if (targetKey != null) {
nodeLabels.putIfAbsent(targetKey, anyValueToString(it.getValue()));
}
}
if (!nodeLabels.containsKey(NODE_IDENTIFIER_HOST_NAME)) {
for (final var it : attributes) {
if (SERVER_ADDRESS.equals(it.getKey())) {
nodeLabels.put(NODE_IDENTIFIER_HOST_NAME, anyValueToString(it.getValue()));
break;
}
}
}
return nodeLabels;
}

private static Map<String, String> buildLabels(List<KeyValue> kvs) {
return kvs
.stream()
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@

package org.apache.skywalking.oap.server.receiver.otel.otlp;

import io.opentelemetry.proto.common.v1.AnyValue;
import io.opentelemetry.proto.common.v1.KeyValue;
import io.opentelemetry.proto.metrics.v1.ExponentialHistogram;
import io.opentelemetry.proto.metrics.v1.ExponentialHistogramDataPoint;
import io.opentelemetry.proto.metrics.v1.Metric;
Expand Down Expand Up @@ -131,4 +133,33 @@ public void testAdaptExponentialHistogram() throws NoSuchMethodException, Invoca
assertTrue(histogramMetric.getBuckets().containsKey(-Math.pow(base, 17)));
assertEquals(2, histogramMetric.getBuckets().get(-Math.pow(base, 17)));
}

private static KeyValue attribute(final String key, final String value) {
return KeyValue.newBuilder().setKey(key).setValue(AnyValue.newBuilder().setStringValue(value)).build();
}

// OTel Collector v0.127.0 and later send a Prometheus target's host as server.address only.
@Test
public void testHostNameFromServerAddress() {
final Map<String, String> labels = OpenTelemetryMetricRequestProcessor.nodeLabels(List.of(
attribute("service.name", "windows-monitoring"),
attribute("server.address", "172.25.0.1"),
attribute("service.instance.id", "172.25.0.1:9182")
));
assertEquals("172.25.0.1", labels.get("node_identifier_host_name"));
assertEquals("windows-monitoring", labels.get("job_name"));
assertEquals("172.25.0.1", labels.get("server_address"));
}

@Test
public void testLegacyHostAttributesWinOverServerAddress() {
assertEquals("win-host", OpenTelemetryMetricRequestProcessor.nodeLabels(List.of(
attribute("server.address", "172.25.0.1"),
attribute("host.name", "win-host")
)).get("node_identifier_host_name"));
assertEquals("10.211.55.3", OpenTelemetryMetricRequestProcessor.nodeLabels(List.of(
attribute("net.host.name", "10.211.55.3"),
attribute("server.address", "10.211.55.3")
)).get("node_identifier_host_name"));
}
}
Loading
Loading