Wayseer

User guideWayseer 0.28.3Contents

Prometheus

The prometheus module shows what a Prometheus server scrapes: each scrape target, grouped by job, on the hosts named by their instance labels. Every metric the server describes can be charted. With an Alertmanager, firing alerts show on the entities they are about. It reads through the HTTP API, and changes nothing unless you allow it to silence alerts.

Configuration

modules:
  - kind: prometheus
    name: prom
    options:
      url: http://localhost:9090

url defaults to http://localhost:9090. Put no credentials in the URL. A server behind basic or bearer auth takes them from a file or an environment variable:

modules:
  - kind: prometheus
    name: prom
    options:
      url: https://prometheus.example.com
      auth: bearer
      secret_file: ~/.config/wayseer/prom-token   # or secret_env: PROM_TOKEN
      alertmanager:
        url: https://alertmanager.example.com
        auth: basic
        username: wayseer
        secret_env: ALERTMANAGER_PASSWORD

secret_keyring: <service>/<account> reads the secret from the system keyring instead (secrets). The secret is read once, when the module starts, and sent only in the Authorization header. It never shows in the log, in an error or on screen. Alertmanager has its own credentials; the server's are never sent to it.

The options below go under options.

url

The server's address, http:// or https://.

Type
string

auth

How to authenticate: none, basic or bearer.

Type
string
Default
none

username

The user, for basic.

Type
string

secret_file

A file holding the secret; ~/ is the home directory.

Type
string

secret_env

Or the environment variable holding it.

Type
string

secret_keyring

Or the keyring entry holding it, as service/account.

Type
string

timeout

Longest wait for one request, 100ms to 5m.

Type
duration
Default
10s

interval

How often targets are read, 1s to 1h.

Type
duration
Default
30s

kinds

Job name to entity kind; other jobs' targets are services.

Type
map of string
Default
node and node-exporter are host

alertmanager

An Alertmanager whose firing alerts show on their entities, with its own credentials.

Type
mapping

alertmanager.url

The server's address, http:// or https://.

Type
string

alertmanager.auth

How to authenticate: none, basic or bearer.

Type
string
Default
none

alertmanager.username

The user, for basic.

Type
string

alertmanager.secret_file

A file holding the secret; ~/ is the home directory.

Type
string

alertmanager.secret_env

Or the environment variable holding it.

Type
string

alertmanager.secret_keyring

Or the keyring entry holding it, as service/account.

Type
string

flows

Queries whose answers are traffic between entities in the world.

Type
list

flows[].query

PromQL giving a rate per second for each source and destination.

Type
string
Default
required

flows[].from

Where traffic comes from.

Type
mapping
Default
required

flows[].from.label

The label whose value names the entity.

Type
string
Default
required

flows[].from.kind

The entity's kind.

Type
string
Default
service

flows[].from.make

Make an entity of kind for a value no module found; needs kind.

Type
bool

flows[].to

Where traffic goes.

Type
mapping
Default
required

flows[].to.label

The label whose value names the entity.

Type
string
Default
required

flows[].to.kind

The entity's kind.

Type
string
Default
service

flows[].to.make

Make an entity of kind for a value no module found; needs kind.

Type
bool

flows[].unit

What the rate counts: requests, bytes or messages.

Type
string
Default
required

max_made

Most entities flow ends with make may make, 1 to 10000.

Type
integer
Default
1000

series

Queries whose answers are a metric of entities in the world.

Type
list

series[].query

PromQL giving a series for each entity.

Type
string
Default
required

series[].metric

The metric's name, such as queue.depth.

Type
string
Default
required

series[].unit

Its unit, such as count, bytes or percent.

Type
string
Default
none, a plain number

series[].description

What it measures, shown with it.

Type
string

series[].entity

The entity each series is of.

Type
mapping
Default
required

series[].entity.labels

Such as [namespace, pod]; the last alone is tried too.

Type
list of string
Default
required

series[].entity.kind

The entity's kind.

Type
string
Default
service

What shows

service (the server)

ok

Attributes
url

prometheus/job

warn when some of its targets are down, crit when all are

Attributes
targets

service (a target)

up, down with the scrape error, or unknown until first scraped

Attributes
job, instance, scrape_url, scrape_interval, scrape_duration (seconds)

alert

by severity while it notifies, as below; unknown while silenced or inhibited, saying which

Attributes
state (firing, silenced, inhibited or muted), summary, and each label as label.<name>

host

as a target, when it is one; otherwise ok while any target on it is up, and down when none answers

Attributes
as for a target, when it is one

any kind a flow end makes

unknown

Attributes
flow_label

With flows, entities that other modules found also talk to each other, with a rate (flows).

A target is a member of its job, and a job is a member of the server. A target runs on the host its instance label names, without the port.

A node exporter target is the host itself, named by the exporter's nodename, and so is every target of a job that kinds maps to host (node and node-exporter by default). kinds can map any other job to a kind, such as {postgres: database}.

Metrics

Every gauge and counter the server describes can be charted by its own name, such as http_requests_total. Counters are shown as a rate per second. To see one across every target, run it in Grid, and Tab completes the name (lenses):

>grid.show metric=http_requests_total kind=service

Grid shows the hottest it has room for, at most 500. The module asks the server for them with topk at the end of the time window, then fetches only their series, so a Grid over thousands of targets stays quick.

Hosts and services also answer the metric names every module shares, so a Grid of cpu.utilisation puts Prometheus hosts beside hosts from other sources:

MetricForFrom
cpu.utilisationhosts, servicesnode_cpu_seconds_total
process_cpu_seconds_total
memory.utilisationhostsnode_memory_MemAvailable_bytes
node_memory_MemTotal_bytes
memory.rssservicesprocess_resident_memory_bytes
disk.read
disk.write
hostsnode_disk_read_bytes_total
node_disk_written_bytes_total
net.receive
net.transmit
hostsnode_network_*_bytes_total, without lo

With series, queries you write become metrics of entities other modules found, such as pods (series).

Flows

flows turns queries you write into traffic between entities, for the Flow lens. Each series a query answers is a rate per second from the entity one label names to the entity another names. With a service mesh such as Istio, that joins the deployments the Kubernetes module found:

modules:
  - kind: prometheus
    name: prom
    options:
      url: http://localhost:9090
      flows:
        - query: sum by (source_workload, destination_workload) (rate(istio_requests_total{reporter="destination"}[1m]))
          from: {label: source_workload, kind: k8s/deployment}
          to: {label: destination_workload, kind: k8s/deployment}
          unit: requests

Any query with a label for each end works. This one reads bytes between hosts from an exporter of your own:

      flows:
        - query: sum by (src, dst) (rate(net_sent_bytes_total[1m]))
          from: {label: src, kind: host}
          to: {label: dst, kind: host}
          unit: bytes

The flows[] rows of the options table above give each field. A label's value names the entity of that kind with that name or ID, or that the identity rules match, so db-07:5432 names the host db-07. Series between the same two entities add up, so a query need not sum away labels such as response_code. Traffic from an entity to itself is left out. Where two queries give the same two entities, the first query's rate is used.

A value that names no entity in the world adds nothing unless its end has make: true. The module's health note counts such series, and those whose value names several entities that are not the same, such as checkout in two namespaces. Each is matched again at the next read, so a flow appears once the module that finds its ends has read them.

Ends that make entities

With make: true, an end whose value names no entity makes one of its kind, named by the value, so traffic stands on its own where no other module finds its ends, as with a tracing service graph:

      flows:
        - query: sum by (client, server) (rate(traces_service_graph_request_total[1m]))
          from: {label: client, kind: service, make: true}
          to: {label: server, kind: service, make: true}
          unit: requests

make needs a kind. A value is still matched first: when another module found an entity of that kind and name, or this one scrapes it, the flow joins that one and nothing is made, and a made entity gives way to one found later. A made entity has status unknown and the attribute flow_label, the label that named it. It goes when no series names it, except that a query that fails keeps its last traffic and the entities that traffic joins.

At most max_made entities are made, 1000 unless set, from 1 to 10000. Ends past that add no traffic, and the health note says how many were left out.

The queries run at interval. If one fails, the health note says so, without the query or what the server said about it, and its last traffic stays. The queries are yours: the language model never writes or sees them.

Service graphs

Tempo's metrics generator and the OpenTelemetry Collector's servicegraph connector turn traces into metrics of which service calls which. With one Prometheus instance and nothing else, this config shows those services talking in Flow and Topology, with each service's failed requests and latency:

modules:
  - kind: prometheus
    name: prom
    options:
      url: http://localhost:9090
      flows:
        - query: sum by (client, server) (rate(traces_service_graph_request_total[1m]))
          from: {label: client, kind: service, make: true}
          to: {label: server, kind: service, make: true}
          unit: requests
      series:
        - query: sum by (server) (rate(traces_service_graph_request_failed_total[5m]))
          metric: requests.failed
          unit: per_second
          description: failed requests the service served
          entity: {labels: [server], kind: service}
        - query: histogram_quantile(0.95, sum by (server, le) (rate(traces_service_graph_request_server_seconds_bucket[5m])))
          metric: latency.p95
          unit: seconds
          description: 95th percentile time to serve a request
          entity: {labels: [server], kind: service}

Both write the same metric names. If yours differ, such as with a prefix added on export, change them in each query. What it shows:

  • Each service named in client or server is a service with status unknown (ends that make entities), unless another module found it, as the Kubernetes module finds services, when the flow joins that one.
  • Flow draws a band for each caller and callee, as wide as its requests per second. Topology places every service, joined to those it calls, though no host is under them.
  • Grid, Detail and Timeline show requests.failed and latency.p95 for each service that is called. A service that only calls others, such as a user's browser, has neither.

It can't show the spans themselves, or a call between two services that no trace saw.

Series

series turns queries you write into a metric of entities that other modules found, such as the pods the Kubernetes module or a file found. Each series a query answers belongs to the entity its labels name:

modules:
  - kind: prometheus
    name: prom
    options:
      url: http://localhost:9090
      series:
        - query: max by (namespace, pod, queue) (shop_queue_depth)
          metric: queue.depth
          unit: count
          description: messages waiting in the pod's queues
          entity: {labels: [namespace, pod], kind: pod}

The values of entity.labels, joined by /, name the entity by its ID, so shop and checkout-6d8f9c7b5-x2k4p name the pod shop/checkout-6d8f9c7b5-x2k4p. If no entity has that ID, the last label's value names it by its name, or as the identity rules match. Series naming the same entity add up, so the query above gives each pod the depth of all its queues.

The metric joins the catalog under its name: Grid, Detail and Timeline show it for every entity of that kind, and the language model can ask for it, seeing its name, unit and description. An entity's own module wins: if it offers a metric of the same name, that one is shown. unit is one of bytes, bytes_per_second, bits, bits_per_second, percent, ratio, seconds, count or per_second; without it the value is a plain number.

A query runs when its metric is shown, over the time shown, as written: Wayseer adds no selector to it. A series that names no entity in the world is left out, and the module's health note counts such series, and those naming several entities, by metric. If the server refuses a query, the error names the metric and why, such as bad_data, without the query or what the server said about it. The queries are yours: the language model never writes or sees them.

Recipes give kinds, series and flows for common exporters, such as MySQL's, Redis's, MongoDB's, Kafka's and Netdata's.

Alerts

With alertmanager, each firing alert is an entity of kind alert, named by its alertname. It is a member of the entity its labels name: the target its job and instance name, else the host its instance names, else its job, else the server. Resolved alerts leave the world. On the entity they are about:

  • A critical alert makes the entity crit. Any other severity but info or none makes it warn. An alert never lowers a status the entity already has.
  • The alert names show as the entity's reason, and in its alerts attribute.
  • Silenced or inhibited alerts are listed in alerts_muted, and leave the status alone.
  • Firing, being muted and resolving are events on the Timeline.

If Alertmanager cannot be read, the last alerts stay and the module's health says why.

Actions

With alertmanager, the module offers two actions on alerts. Neither runs unless the instance's actions lists it, and each waits for you to confirm it.

silence

Makes a silence matching every one of the alert's labels exactly, for for: from 15 minutes to 7 days, 1 hour if not given. An alert already silenced is refused.

unsilence

Expires the silences on the alert that this instance made. Silences made elsewhere are left alone, and are never edited.

modules:
  - kind: prometheus
    name: prom
    actions: [silence, unsilence]
    options:
      url: http://localhost:9090
      alertmanager:
        url: http://localhost:9093

A silence's createdBy is the instance's name (prom above) and its comment is Wayseer; that is how unsilence knows its own. Neither carries your name or email. The alert shows as silenced from the module's next read of Alertmanager, within interval.

If Alertmanager refuses the credentials, the action's error says so and nothing more. Any other error is cut to its first line, and never holds the secret.