Logs
Look up every field a Bot Manager report log carries, the classification values it records, and where each log surfaces.
Bot Manager writes a report log. Each line is one JSON object naming the request, the score the function calculated for it, the rules the request matched, and the action the function applied. The internal_logs argument sets which requests the function writes a line for, and log_tag sets the tag that identifies the instance the line came from. For more information, refer to Arguments.
The two editions do not write the same payload. Bot Manager Lite writes 14 fields, and Bot Manager documents 25. A line read from a Lite instance is not a full-edition line with fields missing.
This page lists the fields each edition writes, the values classified and bot_category take, what the under-evaluation verdict means, where each log surfaces, and how long each dataset is retained.
Fields
Every report line opens with the same prefix, then a single JSON object. The prefix is [Bot-Protection][<log_tag>] Report:, where the second bracket carries the value of the instance’s log_tag argument. The host is not in the prefix. It is a field inside the object, so the prefix identifies the instance that wrote the line and nothing else about the request.
Bot Manager Lite fields
A Bot Manager Lite line, from an instance tagged storefront-bots:
Bot Manager Lite writes these 14 fields, and it writes all 14 on every line:
| Field | Type | What it carries |
|---|---|---|
action | string | The action the function applied. It is allow whenever the score does not reach the threshold and the function identifies no attack, and otherwise the value the action argument carries |
asn | string | The ASN the request came through |
bot_category | string | The categories of the rules the request matched, joined by commas |
classified | string | The verdict the function reached. The four values are listed in Classification |
fingerprint | string | The identifier Azion derives for the client behind the request, as four underscore-separated segments |
geoip_country | string | The code of the country the request came from |
geoip_region | string | The code of the region the request came from |
host | string | The host the request was made to |
http_user_agent | string | The User-Agent header the request sent, empty when the request sent none |
matched_rules | array of numbers | The IDs of the rules the request matched |
remote_addr | string | The IP address that initiated the request |
request_id | string | The unique identifier of the request |
request_uri | string | The request URI |
score | number | The score the function calculated for the request |
Bot Manager Lite carries no log_tag field, because the prefix already names the instance. It also carries no header array, no CAPTCHA field, and no cookie field. The score is the sum of the increments of the rules in matched_rules, which the rule table publishes one by one. For more information, refer to Bot Manager Lite.
Bot Manager fields
Bot Manager documents 25 fields. Ten of them, and the type of two more, are the difference between the two editions:
| Field | Type | What it carries |
|---|---|---|
action | string | The action the function applied. It is allow whenever the score does not reach the threshold and the function identifies no attack, and otherwise the value the action argument carries |
asn | string | The ASN the request came through |
azion_fingerprint | string | The fingerprint Azion identified for the request |
bot_category | string | The category the request best fits. The values are listed in Classification |
bot_characteristics | array of strings | Every static-rule violation category and every attack the request matched |
bot_mode | string | The mode the function scored the request in |
bytes_sent | number | The Content-Length of the request |
challenge_solved | boolean | Whether the client solved a CAPTCHA challenge. It reports a value only where a CAPTCHA function runs alongside Bot Manager |
classified | string | The verdict the function reached. The four values are listed in Classification |
disabled_matched_rules | array of numbers | The IDs of the disabled static rules the request matched. A disabled rule that matches adds nothing to the score |
engine_version | number | The engine that identified the client. It is present only where the dynamic rules are turned on |
geoip_country | string | The code of the country the request came from |
geoip_region | string | The name of the region the request came from |
host | string | The host the request was made to |
http_user_agent | string | The User-Agent header the request sent |
log_tag | string | The tag identifying the instance the line came from, set by the log_tag argument |
matched_rules | array of numbers | The IDs of the enabled static rules the request matched |
persist_cookies | boolean | Whether the function finds the client keeping the bot protection cookies |
remote_addr | string | The IP address that initiated the request |
request_headers | array of strings | Every header named in the log_headers argument that is not on the forbidden list, written as key value with the value base64-encoded |
request_id | string | The unique identifier of the request |
request_method | string | The request method |
request_uri | string | The request URI |
score | number | The score the function calculated for the request |
sent_http_content_type | string | The Content-Type header the request sent |
One further key, concat_headers, appears in a Bot Manager report object without a published definition. Its value is a comma-separated list of header names.
request_headers is bounded. Each key and value deducts from a total of 10,000 characters, and once that total is reached no further header is written into the array. A header value is base64-encoded, so bnVsbA== in the array is the encoding of null, which is what a header the request did not send decodes to. For more information, refer to Firewall limits.
Two fields point at configuration rather than at the request. An instance that sets no log_tag is tagged with the request host, so two instances that both keep the shipped default are indistinguishable in the log. The fingerprint is the identifier Azion derives for a client, which makes it the field that groups the repeat requests of one device. Bot Manager consolidates fingerprint data across requests and the GraphQL API returns it, so you can act on a device or user identified as a bad bot at a threshold rather than on a single request.
Classification
Two fields carry the verdict, and they are computed differently. classified compares the score against the threshold in force. bot_category names the categories of the rules the request matched, whatever the verdict.
classified is therefore relative, not a property of the request on its own. The same score, from the same rules, is classified two ways under two thresholds:
score | matched_rules | threshold | classified | action |
|---|---|---|---|---|
| 28 | [1,10,18,19,20] | 30 | legitimate | allow |
| 28 | [1,10,18,19,20] | 1 | bad bot | deny |
| 28 | [26,10,18,19,20] | 30 | legitimate | allow |
Raising a threshold relabels the traffic as well as stopping the action from firing. A classification count is comparable across two periods only when the threshold was the same in both, so record the threshold alongside any figure you take from the classification charts.
bot_category is derived from the rules the request matched, and it is a comma-joined list rather than one value. A request that matched rules from two categories carries both, as Bad Bot Signatures, Malicious Intent detected. Because the categories follow the rules and the verdict follows the threshold, a line classified legitimate can still name bad-bot categories: the categories say which rules fired, and classified says whether their total crossed the threshold.
The pairs the two fields take, and how a request earns each one, are below. Real-Time Metrics groups its charts by the same pairs.
classified | bot_category | How the request is identified |
|---|---|---|
| Good Bot | Good Bot | By user agents associated with social networks, content aggregators, monitoring bots, and search engines |
| Bad Bot | Bad Bot Signatures | By user agents known for malicious behavior, including malicious signatures and missing or anomalous headers |
| Bad Bot | Scripted Bots | By user agents that indicate automation, such as headless or dalvik, and by unusual user agent length |
| Bad Bot | Malicious Browser Behavior | By missing or forged essential cookies, missing required HTTP headers, and cookie validation failures |
| Bad Bot | Malicious Intent Detected | By unusual HTTP headers and methods, such as TRACE |
| Bad Bot | Reputation Intelligence | By checking the request IP address against known reputation lists |
| Bad Bot | Brute Force | By a high frequency of login attempts, IP address variation, and error patterns |
| Bad Bot | Scraping | By high URL access variability and request frequency |
| Bad Bot | Crawling | By URL variation patterns and the request frequency of a content crawler |
| Bad Bot | Credential Stuffing | By login attempt frequency, error patterns, and attempts on multiple accounts |
| Bad Bot | Credential Cracking | By request frequency and specific error patterns |
| Bad Bot | Account Takeover | By anomalous request patterns and high geographic variation |
| Legitimate | Non-Bot Like | No suspicious behavior and no bot pattern was identified |
| Under Evaluation | Under Evaluation | There is not enough data for a complete classification |
The log writes the four classified values in lower case, as bad bot, good bot, legitimate, and under evaluation. The charts built on them title-case the same values, so a value read from a chart and a value read from a line differ in case and in nothing else.
Each value has a condition. legitimate means the function did not identify a bot and had enough fingerprint data to rule out an attack. good bot means the function identified no attack and matched the user agent against a good user agent pattern. bad bot means the request reached the score threshold, or the function identified it as an attack. under evaluation means the function did not identify a bot and did not have enough fingerprint data to rule out an attack.
For a request classified good bot, the category is the good-bot type its user agent matched: Aggregator Bot, Enterprise Bot, Monitoring Bot, Search Engine Bot, or Social Network Bot.
Under evaluation
under evaluation is the verdict the function records when it has not identified a bot and does not yet hold enough fingerprint data to rule out an attack. Both classified and bot_category carry it, and it is one of the four values the traffic charts split requests across.
A fingerprint stays under evaluation until Bot Manager has consolidated enough data about it, which is why a fingerprint the function is seeing for the first time carries the verdict. For more information, refer to Bot scoring.
The verdict is an absence of evidence, not a finding. Counting an under-evaluation line as legitimate traffic overstates your clean traffic, and counting it as bot traffic overstates the opposite. Read the three decided verdicts against each other and read this one as the share of traffic the function has not yet placed.
Where logs surface
Four surfaces carry Bot Manager data, and they do not carry the same thing. Real-Time Events and Data Stream carry the report line itself, field for field. Real-Time Metrics and the Real-Time Metrics GraphQL API carry counts aggregated from those fields, so a single request is not retrievable from either one.
In Real-Time Events, a report line is a record of the functionConsoleEvents dataset, queried at https://api.azion.com/v4/events/graphql:
The query answers 200 and returns one record per line the function wrote. line holds the prefix and the JSON object together, level is LOG, and lineSource is CONSOLE. functionId is the id of the installed function, and configurationId is the id of the workload the request arrived on, not the id of the firewall the instance runs on. Because every line names its instance in the prefix, the log_tag value is what separates one instance’s lines from another’s. In Azion Console the same records are rendered with a Time column and a Log Body column, and the interface carries no Bot Manager field of its own: the JSON object arrives whole rather than split into columns you can sort. For more information, refer to Real-Time Events.
Azion CLI does not return these lines. While an instance scores live traffic, azion logs cells --function-id <function-id> and azion logs http print nothing, filtered or not, as the dataset returns every line the function writes. An empty CLI result therefore does not show a silent function.
Data Stream forwards the report line to an endpoint you configure, reading it from the Functions data source, which requires a subscription to Functions. The forwarding is real time, so a dashboard or an alert built on that endpoint sees a line as the function writes it. The fields an endpoint receives depend on its type. For more information, refer to Endpoints.
Real-Time Metrics carries a Bot Manager page with an Overview dashboard and a Breakdown dashboard. Its charts aggregate the classification the log carries: Top Bot Action, for one, groups requests by the action the function applied. Metrics are generated almost in real time, with an aggregation interval of up to 60 seconds. For more information, refer to Real-Time Metrics.
Two GraphQL datasets carry the same aggregated data for querying. botManagerMetrics groups by the classification, the action, the mode, the CAPTCHA result, the host, and the geography of a request. botManagerBreakdownMetrics groups by the URLs bot traffic reached and the IP addresses it came from. The fields of each are listed on the Real-Time Metrics GraphQL API Fields page, under botManagerMetrics and botManagerBreakdownMetrics.
Retention
The two Real-Time Metrics GraphQL datasets are retained for different periods, because they hold different data. botManagerMetrics is retained for 2 years, which is also how long the charts built on it keep their data. botManagerBreakdownMetrics is retained for 60 days. A query that reaches further back than 60 days therefore returns the classification counts and not the URLs and IP addresses behind them.
Data Stream changes the question rather than answering it. It forwards a copy of the report line to an endpoint you own, so how long that copy survives is a property of the endpoint and of what you configured it to keep, not of Bot Manager.