Log Parsing Pipeline :: Kloudfuse Docs
Log Parsing Pipeline
Overview
Kloudfuse ingests log lines through the ingester service, which writes the incoming log payloads to a Kafka topic. Each payload is then read by the Logs Parser service, and is unmarshalled (unwrapped) into a JSON string. This JSON string is the input into the log parsing pipeline.
- Unwrap/unmarshall converts the incoming Log Event Payload into a JSON string, to prepare it for the Remap stage and subsequent processing.
- Δ indicates that the fields changed at the previous stage of the pipeline.
- Wrap/marshall converts the processed log into a protoBuf load.
| Kloudfuse truncates log messages larger than 64KB. See Pre-parse. |
Remap
Remap is the first stage in the log parsing pipeline. It maps fields from the incoming payload to a field in the internal representation of the log event. If the log payload is in msgpack format or proto format, Parse Builder converts them into a JSON format before beginning a remap.
We support logs ingestion from many agents into various formats (JSON, msgpack, and proto).
Remap extracts the following fields from the log payload:
- Log message
- Timestamp (see Timestamp Handling for how Kloudfuse selects and validates timestamps)
- Log source
- Labels and Tags
- Log Facets
This stage is fully configurable; you can instruct how to map fields from the incoming log payload to a field in Parse Builder.
The code snippet in the Sample remap function is an exhaustive list, and covers all possible arguments of the remap function. Most of these fields ship with a good set of defaults, so you typically don’t have to define all arguments and override them, unless you plan to deviate from default configuration. Note that all the fields in the Remap function must be specified in JSONPath notation.
Sample remap function
- remap:
args:
kf_source:
- "$..logSource"
kf_msg:
- "$..logMessage"
kf_timestamp:
- "$..myTimestamp"
conditions:
- matcher: "__kf_agent"
value: "fluent-bit"
op: "=="
Relabel
The Relabeling stage of the log parsing pipeline operates on labels and tags extracted from the previous stage, Remap.
In this stage you can make the following changes:
- Add, Drop, or Replace a label.
- Create a new label by combining various label value(s).
- Keep or Drop a log event that matches label value(s).
Here are a few sample config for relabel function:
Drop labels where values don’t match kata@webserver.* regex
- relabel:
args:
- action: "keep"
- sourceLabels: "subsystem,server"
- regex: "kata@webserver.*"
- separator: "@"
Add a label env with value "production"
- relabel:
args:
- action: "replace"
- replacement: "production"
- targetLabel: "env"
PreParsing
Preparsing stage is internal to the Kloudfuse stack; it is not configurable. Preparsing truncates log messages that exceed 64KB.
Grammar
Kloudfuse automatically detects log facets and timestamp from the input log line. This is a heuristic-based approach, and the extracted facets may not be accurate all the time. Users can therefore define custom grammars to facilitate log facet extraction.
Dissect patterns
The dissect patterns approach uses a simple text-based tokenizer as a set of fields and delimiters that describe the textual format.
Grok patterns
Grok patterns are based on regexes (named regexes).
Sample grammar config with dissect and grok patterns
kf_parsing_config:
config: |-
- parser:
dissect:
args:
- tokenizer: '%{timestamp} %{level} [LLRealtimeSegmentDataManager_%{segment_name}]'
- parser:
grok:
args:
patterns:
- (%{NGINX_HOST} )?"?(?:%{NGINX_ADDRESS_LIST:nginx_ingress_controller_remote_ip_list}|%{NOTSPACE:source_address})
Kf-Parse
The KF-Parse stage of the log parsing pipeline automatically detects facets and generates fingerprints. It is also internal, and not configurable.
Transform
Transform is the last function of the pipeline, where you can derive labels based on extracted facets from the previous stages.
Add value eventSource to label source
- transform:
args:
- action: "facet_to_label_map"
- sourceLabels: "@eventSource"
- targetLabel: "source"
conditions:
- matcher: "#source"
op: "=="
value: "awsLogSource"
Write Kafka topic
As the log line completes its journey through the entire log parsing pipeline, the Parse Builder extracts all facets and labels. Kloudfuse then wraps (marshals) all fields into a protoBuf object, and writes it to a Kafka topic, which Pinot subsequently consumes.