Sign inSign up

telemetryflow/telemetryflow-agent

By telemetryflow

•Updated 4 days ago

TelemetryFlow Agent (OTEL Agent)

Image
Developer tools
Monitoring & observability
0

10K+

telemetryflow/telemetryflow-agent repository overview

TelemetryFlow Logo
⁠TelemetryFlow Agent (OTEL Agent)

Version License Go Version OTEL SDK OpenTelemetry


Enterprise-grade telemetry collection agent built on OpenTelemetry Go SDK v1.43.0. Provides comprehensive system monitoring with metrics collection, heartbeat monitoring, and OTLP telemetry export for the TelemetryFlow Platform.

This agent works as the client-side counterpart to the TelemetryFlow Backend Agent Module (NestJS), providing:

  • Agent registration & lifecycle management
  • Heartbeat & health monitoring
  • System metrics collection
  • OTLP telemetry export

⁠TelemetryFlow Ecosystem

TFO-Agent is fully aligned with the TelemetryFlow ecosystem, sharing the same OpenTelemetry SDK version:

graph LR
    subgraph "TelemetryFlow Ecosystem v1.3.5"
        subgraph "Instrumentation"
            SDK[TFO-Go-SDK<br/>OTEL SDK v1.43.0]
        end

        subgraph "Collection"
            AGENT[TFO-Agent<br/>OTEL SDK v1.43.0]
        end

        subgraph "Processing"
            COLLECTOR[TFO-Collector<br/>OTEL v0.146.1]
        end
    end

    APP[Application] --> SDK
    SDK -->|OTLP| AGENT
    HOST[Host Metrics] --> AGENT
    AGENT -->|OTLP gRPC/HTTP| COLLECTOR
    COLLECTOR --> BACKEND[TelemetryFlow<br/>Platform]

    style SDK fill:#81C784,stroke:#388E3C
    style AGENT fill:#64B5F6,stroke:#1976D2
    style COLLECTOR fill:#FFB74D,stroke:#F57C00
ComponentVersionOTEL BaseDescription
TFO-Agentv1.2.0SDK v1.43.0Telemetry collection agent
TFO-Go-SDKv1.2.0SDK v1.43.0Go instrumentation SDK
TFO-Collectorv1.1.8Collector v0.147.0Central telemetry collector

⁠Features

⁠OpenTelemetry Core
  • OpenTelemetry SDK v1.43.0: Built on standard OTEL Go SDK (aligned with TFO-Go-SDK)
  • OTLP Export: OpenTelemetry Protocol for metrics, logs, and traces
  • Multi-Signal Support: Metrics, logs, and traces collection
⁠Agent Lifecycle (Backend Integration)
  • Agent Registration: Auto-register with TelemetryFlow backend
  • Heartbeat Monitoring: Regular health checks to backend
  • Health Status Sync: Report agent health and system info
  • Activation/Deactivation: Remote agent control from backend
⁠System Monitoring
  • System Metrics Collection: CPU, memory, disk, and network metrics
  • Docker Container Monitoring: Per-container CPU, memory, network, disk I/O, and PID metrics via Docker Engine API
  • cAdvisor Metrics Scraping: Prometheus endpoint scraper for cAdvisor container metrics
  • Process Monitoring: Track running processes
  • Resource Detection: Auto-detect host, OS, and container info
⁠Reliability
  • Disk-Backed Buffer: Resilient retry buffer for offline scenarios
  • Auto-Reconnection: Automatic retry with exponential backoff
  • Graceful Shutdown: Signal handling (SIGINT, SIGTERM, SIGHUP)
⁠Platform
  • Cross-Platform: Linux, macOS, and Windows support
  • LEGO Building Blocks: Modular architecture for easy extensibility

⁠Quick Start

šŸš€ New to TFO-Agent? Check the Quick Start Guide⁠ for step-by-step setup with Docker, Kubernetes, or binary installation.

⁠From Source
# Clone the repository
git clone https://github.com/telemetryflow/telemetryflow-agent.git
cd telemetryflow-agent

# Build
make build

# Run
./build/tfo-agent --help
⁠Docker
# Copy environment template
cp .env.example .env

# Edit .env with your configuration
vim .env

# Build and start
docker-compose up -d --build

# View logs
docker-compose logs -f tfo-agent

# Stop
docker-compose down
⁠Using Docker Directly
# Build image
docker build \
  --build-arg VERSION=1.1.8 \
  --build-arg GIT_COMMIT=$(git rev-parse --short HEAD) \
  --build-arg GIT_BRANCH=$(git rev-parse --abbrev-ref HEAD) \
  --build-arg BUILD_TIME=$(date -u '+%Y-%m-%dT%H:%M:%SZ') \
  -t telemetryflow/telemetryflow-agent:1.1.8 .

# Run container
docker run -d --name tfo-agent \
  -p 4317:4317 \
  -p 4318:4318 \
  -p 8888:8888 \
  -p 13133:13133 \
  -v /path/to/config.yaml:/etc/tfo-agent/tfo-agent.yaml:ro \
  -v /var/lib/tfo-agent:/var/lib/tfo-agent \
  telemetryflow/telemetryflow-agent:1.1.8
⁠OTEL Collector Ports
PortProtocolDescription
4317gRPCOTLP gRPC (v1 & v2)
4318HTTPOTLP HTTP (v1 & v2)
8888HTTPOTEL Collector metrics
8889HTTPPrometheus exporter
13133HTTPHealth check
55679HTTPzPages (debugging)
1777HTTPpprof (profiling)
⁠OTLP Endpoints (Dual Ingestion)

The TFO-Collector supports both TelemetryFlow (v2) and OTEL Community (v1) endpoints:

TelemetryFlow Platform (Recommended):

POST http://localhost:4318/v2/traces
POST http://localhost:4318/v2/metrics
POST http://localhost:4318/v2/logs

OTEL Community (Backwards Compatible):

POST http://localhost:4318/v1/traces
POST http://localhost:4318/v1/metrics
POST http://localhost:4318/v1/logs

gRPC: localhost:4317 (both v1 and v2)

⁠Configuration

šŸ“‹ Complete Configuration: See configs/tfo-agent.default.yaml⁠ for a full configuration example showing Node Exporter, Kubernetes, and eBPF collectors integrated with TFO Platform. šŸ”— Integration Guide: See TFO Platform Integration Guide⁠ for architecture diagrams, data flow, and production deployment examples.

Create configuration file at /etc/tfo-agent/tfo-agent.yaml:

# TelemetryFlow Platform Configuration (v1.3.5+)
telemetryflow:
  api_key_id: "${TELEMETRYFLOW_API_KEY_ID}"
  api_key_secret: "${TELEMETRYFLOW_API_KEY_SECRET}"
  endpoint: "${TELEMETRYFLOW_ENDPOINT:-localhost:4317}"
  protocol: grpc # grpc or http
  tls:
    enabled: true
    skip_verify: false
  retry:
    enabled: true
    max_attempts: 3
    initial_interval: 1s
    max_interval: 30s

agent:
  name: "TelemetryFlow Agent"
  hostname: "" # Auto-detected if empty
  tags:
    environment: production

heartbeat:
  interval: 60s
  timeout: 10s

collector:
  system:
    enabled: true
    interval: 15s
    cpu: true
    memory: true
    disk: true
    network: true

exporter:
  otlp:
    enabled: true
    batch_size: 100
    flush_interval: 10s
    compression: gzip

buffer:
  enabled: true
  path: "/var/lib/tfo-agent/buffer"
  max_size_mb: 100
⁠Environment Variables
# TelemetryFlow Platform (v1.3.5+)
export TELEMETRYFLOW_ENDPOINT="localhost:4317"
export TELEMETRYFLOW_API_KEY_ID="tfk_your_key_id"
export TELEMETRYFLOW_API_KEY_SECRET="tfs_your_key_secret"
export TELEMETRYFLOW_ENVIRONMENT="production"

# Agent Configuration
export TELEMETRYFLOW_AGENT_ID="your-agent-id"
export TELEMETRYFLOW_AGENT_NAME="my-agent"

# Logging
export TELEMETRYFLOW_LOG_LEVEL="info"

⁠Usage

# Start agent
tfo-agent start

# Start with custom config
tfo-agent start --config /path/to/config.yaml

# Validate configuration
tfo-agent config validate

# Show version
tfo-agent version

⁠Project Structure

tfo-agent/
ā”œā”€ā”€ cmd/tfo-agent/         # CLI entry point
ā”œā”€ā”€ internal/
│   ā”œā”€ā”€ agent/             # Core agent lifecycle
│   ā”œā”€ā”€ buffer/            # Disk-backed retry buffer
│   ā”œā”€ā”€ collector/         # Metric collectors
│   │   ā”œā”€ā”€ aurora/        # Amazon Aurora CloudWatch/PI/RDS collector
│   │   ā”œā”€ā”€ cadvisor/      # cAdvisor Prometheus scraper collector
│   │   ā”œā”€ā”€ clickhouse/    # ClickHouse database collector
│   │   ā”œā”€ā”€ cockroachdb/   # CockroachDB database collector
│   │   ā”œā”€ā”€ docker/        # Docker container metrics collector
│   │   ā”œā”€ā”€ ebpf/          # eBPF kernel-level metrics collector
│   │   ā”œā”€ā”€ kubernetes/    # Kubernetes metrics collector
│   │   ā”œā”€ā”€ mongodb/       # MongoDB database collector
│   │   ā”œā”€ā”€ mssql/         # Microsoft SQL Server collector
│   │   ā”œā”€ā”€ mysql/         # MySQL/MariaDB/Percona collector
│   │   ā”œā”€ā”€ nodeexporter/  # Node Exporter metrics collector
│   │   ā”œā”€ā”€ postgresql/    # PostgreSQL/RDS PostgreSQL collector
│   │   ā”œā”€ā”€ sqlite3/       # SQLite3 database collector
│   │   ā”œā”€ā”€ system/        # System metrics collector
│   │   └── timescaledb/   # TimescaleDB collector
│   ā”œā”€ā”€ config/            # Configuration management
│   ā”œā”€ā”€ exporter/          # OTLP data exporters
│   └── version/           # Version and banner info
ā”œā”€ā”€ pkg/                   # LEGO Building Blocks
│   ā”œā”€ā”€ api/               # HTTP API client
│   ā”œā”€ā”€ banner/            # Startup banner
│   ā”œā”€ā”€ config/            # Config loader utilities
│   └── plugin/            # Plugin registry system
ā”œā”€ā”€ configs/               # Configuration templates
ā”œā”€ā”€ scripts/               # Build/install scripts
ā”œā”€ā”€ build/                 # Build output
ā”œā”€ā”€ Makefile
ā”œā”€ā”€ Dockerfile             # Docker build
ā”œā”€ā”€ docker-compose.yml     # Docker Compose
ā”œā”€ā”€ .env.example           # Environment template
└── README.md

⁠LEGO Building Blocks

The pkg/ directory contains reusable building blocks:

BlockDescription
pkg/bannerASCII art startup banner
pkg/configFlexible configuration loader
pkg/pluginPlugin registry for extensibility
pkg/apiHTTP client for backend communication
⁠Adding Custom Plugins
import "github.com/telemetryflow/telemetryflow/telemetryflow-agent/pkg/plugin"

// Register a custom collector
plugin.Register("my-collector", func() plugin.Plugin {
    return &MyCustomCollector{}
})

// Use the plugin
p, _ := plugin.Get("my-collector")
p.Init(config)
p.Start()

⁠Collected Metrics

MetricTypeDescription
system.cpu.usagegaugeCPU usage percentage
system.cpu.coresgaugeNumber of CPU cores
system.memory.totalgaugeTotal memory (bytes)
system.memory.usedgaugeUsed memory (bytes)
system.memory.usagegaugeMemory usage percentage
system.disk.totalgaugeTotal disk space (bytes)
system.disk.usedgaugeUsed disk space (bytes)
system.disk.usagegaugeDisk usage percentage
system.network.bytes_sentcounterTotal bytes sent
system.network.bytes_recvcounterTotal bytes received
⁠Docker Container Metrics

The Docker collector provides 32 per-container metrics via Docker Engine API:

  • CPU: container.cpu.{usage_percent,usage_total,user,kernel,online_cpus,throttled_periods,throttled_time}
  • Memory: container.memory.{usage,working_set,limit,max_usage,rss,cache,usage_percent}
  • Network: container.network.{rx,tx}_{bytes,packets,errors,dropped} (per-interface)
  • Disk I/O: container.diskio.{read,write}_{bytes,ops}
  • PIDs: container.pids.current
  • State: container.state.{running,stopped,paused,restarting,total}
⁠cAdvisor Metrics

The cAdvisor collector scrapes Prometheus metrics from a running cAdvisor instance:

  • Collects container_* and machine_* metric families
  • Supports counter, gauge, histogram, summary, and untyped metric types
  • Optional metric name allowlist for selective collection
⁠Database Monitoring

The agent provides native collectors for popular databases via direct connection or cloud SDK:

CollectorSourceMetrics
Amazon AuroraAWS SDK (CloudWatch, RDS, Performance Insights)60+ CloudWatch metrics across storage, replication, cache, latency, transactions, availability, backtrack, serverless, global, instance, volume
MySQL/MariaDBDirect connectionGlobal status, InnoDB, replication, Galera, query analytics, schema, MariaDB-specific (Aria, ColumnStore, Spider, query cache, thread pool, user stats)
PostgreSQLDirect connectionpg_stat_activity, pg_stat_database, pg_stat_bgwriter, pg_stat_statements, table stats, replication
MSSQLDirect connectionWait stats, perf counters, index usage, tempdb, agent jobs, query store, file I/O
MongoDBDirect connectionServer status, replica set, sharding, query profiler, collection stats
ClickHouseHTTP APISystem tables, query metrics, merge stats, replication queue
CockroachDBDirect connectionSQL stats, range stats, store metrics, replication
TimescaleDBDirect connectionHypertable stats, chunk stats, compression ratios, continuous aggregates, job health
SQLite3File accessPage cache, WAL metrics, lock contention, integrity checks, table stats
⁠eBPF Metrics (Linux-only)

The eBPF collector provides 28 kernel-level metrics across 7 categories:

  • Syscall: ebpf.syscall.{count,latency_ns,errors} with pid, comm, syscall labels
  • Network: ebpf.tcp.{connections,bytes_sent,bytes_recv,rtt_ns,retransmits}, ebpf.udp.{packets_sent,packets_recv}
  • File I/O: ebpf.fileio.{operations,bytes,latency_ns} with operation label
  • Scheduler: ebpf.sched.{context_switches,runq_latency_ns,oncpu_ns,migrations}
  • Memory: ebpf.memory.{page_faults,major_faults,minor_faults}
  • TCP State: ebpf.tcp.state_transitions with old_state, new_state labels
  • Hubble: hubble.{flows,drops,policy_verdicts,http_requests,dns_queries,l7_errors}

See eBPF Metrics Documentation⁠ for complete catalog.

⁠Development

⁠Prerequisites
  • Go 1.26 or later
  • Make
⁠Build Commands
# Show all commands
make help

# Build Commands
make                # Build agent (default)
make build          # Build agent for current platform
make build-all      # Build agent for all platforms
make build-linux    # Build for Linux (amd64 and arm64)
make build-darwin   # Build for macOS (amd64 and arm64)

# Development Commands
make run            # Build and run agent
make dev            # Run with go run (faster for development)
make lint           # Run linter
make fmt            # Format code
make vet            # Run go vet

# Dependencies
make deps           # Download dependencies
make deps-update    # Update dependencies
make tidy           # Tidy go modules

# Other Commands
make clean          # Clean build artifacts
make install        # Install binary to /usr/local/bin
make uninstall      # Uninstall binary
make docker-build   # Build Docker image
make docker-push    # Push Docker image
make version        # Show version information
⁠Testing
# Run all tests
make test                    # Run unit and integration tests
make test-all                # Run unit, integration, and E2E tests
make test-unit               # Run unit tests only
make test-integration        # Run integration tests only
make test-e2e                # Run E2E tests only

# Run specific tests
make test-run PKG=integrations                      # Run all integration tests
make test-run PKG=domain/agent                      # Run agent domain tests
make test-run TEST=TestPerconaCollector             # Run test by name pattern
make test-run PKG=integrations TEST=TestKafka       # Run specific test in package
make test-list                                       # List available test packages

# Coverage and CI
make test-coverage           # Generate coverage report
make ci-test                 # Run with race detection (CI mode)

# Using test script directly
./scripts/test-specific.sh integrations             # Run all integration tests
./scripts/test-specific.sh -c domain/agent          # Run with coverage
./scripts/test-specific.sh -r TestExporter          # Run with race detector
./scripts/test-specific.sh --ci infrastructure      # CI mode (race + coverage)
./scripts/test-specific.sh -l                       # List available packages
⁠Test Packages
PackageDescriptionTest Files
applicationCLI commands, configuration3
domain/agentAgent lifecycle management2
domain/collector/mysqlMySQL/MariaDB collector3
domain/collector/mongodbMongoDB collector3
domain/collector/postgresqlPostgreSQL collector3
domain/collector/timescaledbTimescaleDB collector1
domain/ebpfeBPF collector4
domain/kubernetesKubernetes collector1
domain/nodeexporterNode Exporter collector1
domain/pluginPlugin registry1
domain/telemetryTelemetry collection2
infrastructure/apiAPI client1
infrastructure/bufferDisk-backed buffer1
infrastructure/configConfiguration loader1
infrastructure/exporterOTLP exporters3
integrations3rd party integrations36
presentation/bannerStartup banner1

⁠Systemd Service

# /etc/systemd/system/tfo-agent.service
[Unit]
Description=TelemetryFlow Agent - CEOP
After=network.target

[Service]
Type=simple
User=telemetryflow
ExecStart=/usr/local/bin/tfo-agent start --config /etc/tfo-agent/tfo-agent.yaml
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable tfo-agent
sudo systemctl start tfo-agent

⁠3rd Party Integrations

TelemetryFlow Agent supports 39+ integrations for enterprise environments across multiple categories.

⁠Integration Categories
CategoryIntegrationsCount
Cloud ProvidersGCP, Azure, Alibaba Cloud, AWS CloudWatch4
InfrastructureProxmox, VMware vSphere, Nutanix, Azure Arc4
Network & IoTCisco (DNA Center/Meraki), SNMP v1/v2c/v3, MQTT3
Kernel/SystemeBPF (syscalls, network, file I/O, scheduler), Cilium Hubble2
APM PlatformsDynatrace, IBM Instana, Datadog, New Relic4
OSS ObservabilitySigNoz, Coroot, HyperDX, OpenObserve, Netdata5
ObservabilityPrometheus, Splunk, Elasticsearch3
Streaming & LogsKafka, Loki, InfluxDB3
TracingJaeger, Zipkin2
Monitoring ToolsTelegraf, Grafana Alloy, Percona PMM, Blackbox, ManageEngine5
CustomWebhook1
⁠Key Differentiators
CapabilityDescription
Unified AgentSingle agent for cloud, infrastructure, network, and system
OTLP-FirstNative OpenTelemetry Protocol support (gRPC & HTTP)
Enterprise ReadyTLS, mTLS, API key authentication, retry with backoff
Hybrid CloudProxmox, VMware, Nutanix, Azure Arc in one agent
Network ObservabilityCisco DNA/Meraki, SNMP v3, MQTT for IoT
Kernel-LeveleBPF for syscalls, network, file I/O, scheduler metrics
ResilientDisk-backed buffer with automatic retry and flush
ExtensiblePlugin architecture for custom integrations

See Integration Documentation⁠ for detailed configuration.

⁠License

Apache License 2.0 - See LICENSE⁠


Copyright (c) 2024-2026 Telemetri Data Indonesia. All rights reserved.

Tag summary

Content type

Image

Digest

sha256:657ff458b…

Size

99.1 MB

Last updated

4 days ago

docker pull telemetryflow/telemetryflow-agent