Monitor and calibrate Bot Manager
Run a Bot Manager instance in observation mode, read the scores in its report log, and set a threshold and rules from your own traffic.
You calibrate a Bot Manager instance by raising its logging, reading the report log it writes, and setting the threshold and the rules from what that log shows.
The same steps apply to Bot Manager and to Bot Manager Lite. The arguments each edition carries differ, and the procedure does not. Firewall best practices covers what each choice costs and how long to hold it; this page covers the steps that produce the choice.
Change one thing at a time. An argument change reaches the request path in about two minutes, so a reading taken sooner describes the configuration that preceded it.
Prerequisites
- A Bot Manager function instance on a firewall, run by a Rules Engine for Firewall rule. For more information, refer to Bot Manager quickstart.
- Access to Azion Console. To sign in, refer to How to access Azion Console.
- A personal token, for the query that reads the report log. For more information, refer to Personal Tokens.
Turn the logging up
An instance writes a report line for the requests internal_logs selects, and the argument ships at 0. Raise it to 2 so that every request produces a line, including one that scores 0, and set action to allow so that a request reaching the threshold is scored and still served.
| Argument | What it does in the observation window |
|---|---|
action | At allow, a request at or above the threshold continues to the application |
internal_logs | At 2, the function writes a line for every request. The value is a number, never a string |
log_tag | The tag the line prefix carries, which is what separates one instance’s lines from another’s |
log_headers | The request headers the function writes into the line. The nine above are the shipped default |
To replace the arguments of an existing instance in Azion Console:
Access Azion Console > Firewalls, then select that firewall.
Select the Bot Manager instance you want to calibrate.
In the Arguments section, enter the object. Bot Manager Lite carries no argument schema, so the section holds a JSON editor and builds no form from it.
The instance scores every request the rule sends it and refuses none of them. Wait about two minutes before reading anything into a response, and leave the window open long enough to cover peak hours, weekly crawlers, and overnight jobs. For how long that is and what it costs, refer to Firewall best practices.
Read the report log in Real-Time Events
Real-Time Events serves each report line as a record of the functionConsoleEvents dataset. In Azion Console those records render in a Time column and a Log Body column, and the interface carries no Bot Manager field of its own, so the JSON object arrives whole rather than split into columns you can sort. Narrow the period to the window you ran, then read the lines whose prefix carries your log_tag.
To read the same records as a query, send the following to https://api.azion.com/v4/events/graphql with an Authorization: Token [TOKEN VALUE] header and a tsRange covering the window:
The query answers 200 and returns one record per line the function wrote. line carries the whole report line, functionId is the id of the installed function, and configurationId is the id of the workload the request arrived on, not the id of the firewall the instance runs on.
A line opens with the prefix [Bot-Protection][<log_tag>] Report: and continues as one JSON object:
Four values carry the calibration. score is the total the function calculated, matched_rules names the rules that produced it, action is what the function applied, and classified is the verdict it reached. In the line above the score is 28 under a threshold of 30, so action reads allow and classified reads legitimate. Under a threshold that 28 reaches, the same request reads bad bot and meets the configured action instead.
bot_category is a comma-joined list of the categories of the rules that matched, so a line classified legitimate can still name bad-bot categories. For every field a line carries and for the four values classified takes, refer to Logs.
The score is on the line and nowhere else. The two Real-Time Metrics GraphQL datasets count requests by classification, action, mode, host, geography, and the URLs bot traffic reached, and neither carries a score, so a score distribution is read from the report log rather than from a query. Those datasets answer the aggregate questions instead: refer to Query Bot Manager data with GraphQL for the classification counts, to Query the top URLs bots reach with GraphQL for the URLs, and to Real-Time Metrics for the same data as charts.
Set the threshold from the scores you read
threshold is the score at which action fires. The value comes from the distribution in your own report log: the clients you recognize cluster at the low end, automated clients cluster higher, and the threshold goes in the gap between the two.
To calibrate it:
Read score on the lines the window produced, and separate the requests you recognize from the ones you do not.
Set it above the scores of the traffic you recognize, and at or below the scores of the traffic you do not.
Replace the instance arguments with the new threshold and the action a request at or above it meets.
A change to an instance reaches the request path in about two minutes.
Read the lines classified legitimate whose score sits just below it. Those requests are one matched rule away from the action.
The instance applies the action to the traffic above the threshold, and the lines below it name the requests that came closest. Repeat the loop: lower the threshold while automated traffic still passes, and raise it when requests you recognize are refused.
Take a rule out of the scoring
A score is the sum of the increments of the rules the request matched, and the line names both: matched_rules carries the IDs and score carries the total. A rule firing on traffic you recognize is found by reading those IDs. No ID belongs in the arguments object before it appears in your own lines.
To stop a rule from raising a score:
Read the lines classified legitimate whose score sits at or near the threshold, and the lines carrying requests you recognize.
matched_rules carries the IDs of the rules each request matched. Compare several lines, so that one request does not settle it.
disabled_rules takes them on Bot Manager Lite, and disabled_static_rules takes them on Bot Manager. Both are arrays.
A rule disabled this way keeps running and adds nothing to the score.
The requests that matched the rule score lower by its increment, and each match lands in the disabled_matched_rules field instead. A rule taken out of the scoring is out of it for every request the instance scores, so the narrower instruments are worth reading first: refer to Firewall best practices. For both arguments and the types they take, refer to Arguments.
Raise the dynamic rules tolerance
Bot Manager documents a dynamic rules method that scores a request against the application’s own traffic baseline. dynamic_rules_tolerance sets how strict that comparison is, and the documented values are soft, medium, and hard, with soft as the documented default.
Move one step at a time. A tolerance change re-scores every request the instance sees, so each step needs an observation window of its own before the next one. Two steps taken together produce their false positives together, with nothing in the lines to say which step produced them.
dynamic_rules_logs_enabled writes the debugging logs of the method. Turn it on for the window and off again when the window closes, because of the volume it adds to the log. For the four arguments the method takes, refer to Arguments.
Keep a copy of the report log
The platform ages the aggregated data out. botManagerMetrics is retained for 2 years and botManagerBreakdownMetrics for 60 days, so a question about the URLs bot traffic reached three months ago has no answer there. A copy in a destination of your own lasts as long as you keep it, and it sits beside the records of the application those requests reached.
Data Stream forwards the report line from the Functions data source to an endpoint you configure, as the function writes it. Lower internal_logs once the observation window closes: at 2 every request produces a line, and that is the volume the stream carries. For the endpoints a stream writes to, refer to Endpoints.