Skip to content
Tech News
← Back to articles

Show HN: Ctrlb-decompose: Strip the noise from logs before sending to LLMs

read original more articles
Why This Matters

Ctrlb-decompose offers a powerful solution for transforming vast amounts of noisy log data into meaningful, structured patterns, significantly reducing complexity and enabling more efficient analysis with LLMs. Its streaming, minimal-memory approach makes it suitable for real-time log processing across various platforms, enhancing debugging and monitoring capabilities for the tech industry and consumers alike.

Key Takeaways

Compress raw log lines into structural patterns with statistics, anomalies, and correlations.

Turn millions of noisy log lines into a handful of actionable patterns — with typed variables, quantile stats, anomaly flags, and severity scoring. Runs as a CLI, in the browser via WASM, or as a Rust library.

$ cat server.log | ctrlb-decompose ┌────────────────────────────────────────────────────────────────────┐ │ ctrlb-decompose: 1,247,831 lines -> 43 patterns (99.9% reduction) │ └────────────────────────────────────────────────────────────────────┘ #1 [ERROR] ██████████████████████ 18,402 (1.5%) <TS> ERROR [<*>] Connection to <ip> timed out after <duration> ip IPv4 unique=12 top: 10.0.1.15 (34%), 10.0.1.22 (28%) duration Duration p50=120ms p99=4.8s #2 [INFO] ████████████████████ 904,221 (72.5%) <TS> INFO [<*>] Request from <ip> completed in <duration> status=<status> ip IPv4 unique=1,847 top: 10.0.1.15 (12%), 10.0.1.22 (8%) duration Duration p50=23ms p99=312ms status Enum unique=3 values: 200 (91%), 404 (6%), 500 (3%)

Website coming soon.

How It Works

ctrlb-decompose uses a two-stage normalization and clustering pipeline that processes logs in a single streaming pass with minimal memory footprint.

┌──────────────────────────────────────────────┐ │ ctrlb-decompose pipeline │ └──────────────────────────────────────────────┘ Raw Log Lines │ ▼ ┌──────────────┐ Strip & parse timestamps (ISO 8601, Apache, │ Timestamp │ syslog, Unix epoch, etc.) into normalized │ Extraction │ <TS> markers with DateTime values. └──────┬───────┘ │ ▼ ┌──────────────┐ Replace integers, floats, IPs, and strings │ CLP │ with compact placeholder bytes. Structurally │ Encoding │ identical lines now produce the same "logtype." └──────┬───────┘ │ ▼ ┌──────────────┐ Tree-based similarity clustering (Drain3) groups │ Drain3 │ logtypes into patterns. Differing tokens become │ Clustering │ <*> wildcards. Incremental — no second pass needed. └──────┬───────┘ │ ▼ ┌──────────────┐ Merge CLP-decoded values with Drain3 wildcard │ Variable │ positions. Classify each variable into semantic │ Extraction │ types: IPv4, UUID, Duration, Enum, Integer, etc. │ & Typing │ └──────┬───────┘ │ ▼ ┌──────────────┐ DDSketch quantiles (p50/p99), HyperLogLog │ Statistics │ cardinality estimation, top-k values, temporal │ Accumulation │ bucketing, and reservoir-sampled example lines. └──────┬───────┘ │ ▼ ┌──────────────┐ Frequency spikes, error cascades, low-cardinality │ Anomaly │ flags, bimodal distributions, and clustered │ Detection │ numeric detection. └──────┬───────┘ │ ▼ ┌──────────────┐ Keyword-based severity (ERROR > WARN > INFO > DEBUG), │ Scoring │ temporal co-occurrence, shared variable correlation, │ & Correlation│ and error cascade detection across patterns. └──────┬───────┘ │ ▼ ┌──────────────┐ │ Output │──── Human (ANSI terminal) / LLM (compact markdown) / JSON └──────────────┘

Stage 1 — CLP Encoding

CLP (Compressed Log Processor) encoding normalizes variable tokens into typed placeholders, so structurally identical lines produce identical logtypes regardless of the actual values:

Input: "Request from 10.0.1.15 completed in 45ms status=200" Logtype: "Request from <dict> completed in <float>ms status=<int>"

... continue reading