Train with IP addresses¶
Pass raw IPv4/IPv6 strings directly to AutoML.fit(). IP detection and numeric
transformation are automatic; no additional configuration is needed.
This example uses synthetic data with labels based on address components. It demonstrates the preprocessing workflow rather than a real-world prediction task.
import numpy as np
import pandas as pd
from supervised import AutoML
X = pd.DataFrame({
"client_ip": [
f"192.0.2.{i}" if i % 2 == 0 else f"2001:db8::{i:x}"
for i in range(1, 129)
],
"request_count": [(i * 7) % 50 + 1 for i in range(1, 129)],
})
y = pd.Series([int(i > 64) for i in range(1, 129)])
X.loc[[9, 89], "client_ip"] = None
print("Training input (first 10 rows, with target):")
print(X.assign(target=y).head(10).fillna("<missing>").to_string(index=False))
automl = AutoML(
results_path="AutoML_IP",
algorithms=["Decision Tree"],
total_time_limit=30,
train_ensemble=False,
stack_models=False,
golden_features=False,
features_selection=False,
kmeans_features=False,
start_random_models=1,
hill_climbing_steps=0,
explain_level=0,
n_jobs=1,
verbose=0,
)
automl.fit(X, y)
new_data = pd.DataFrame({
"client_ip": ["192.0.2.200", "2001:db8::c8", None],
"request_count": [12, 7, 3],
})
print("Prediction input:")
print(new_data.fillna("<missing>").to_string(index=False))
predictions = automl.predict(new_data.copy())
restored = AutoML(results_path="AutoML_IP")
np.testing.assert_array_equal(
predictions, restored.predict(new_data.copy())
)
print(new_data.assign(prediction=predictions))
Both training and prediction use raw strings. The model stores the generated feature schema and reuses it after reload, including for addresses absent during training and missing values that first appear during prediction.
Inspect the input¶
The training input printout begins with:
Training input (first 10 rows, with target):
client_ip request_count target
2001:db8::1 8 0
192.0.2.2 15 0
2001:db8::3 22 0
192.0.2.4 29 0
2001:db8::5 36 0
192.0.2.6 43 0
2001:db8::7 50 0
192.0.2.8 7 0
2001:db8::9 14 0
<missing> 21 0
target is shown alongside the input for inspection; it is passed separately as
y and is not an input feature. <missing> is a display label only: the data
passed to AutoML still contains a missing value.
The raw prediction input is:
Prediction input:
client_ip request_count
192.0.2.200 12
2001:db8::c8 7
<missing> 3
What features are constructed?¶
The input has two columns: client_ip and request_count. AutoML replaces
client_ip with numeric features, keeping only those that vary in training.
There are 14 candidate IP columns. For this training input, six are always
zero and are omitted, leaving 8 IP columns. The nonconstant numeric
request_count column stays unchanged with the Decision Tree configuration used
here, giving 9 model input columns. Components are ordered from left to right
in the address:
| Candidate columns | Maximum count | Meaning |
|---|---|---|
client_ip_IP_Version |
1 | 4 for IPv4, 6 for IPv6, 0 for missing |
client_ip_IP_IPv4_1 through client_ip_IP_IPv4_4 |
4 | Decimal IPv4 octets, each between 0 and 255 |
client_ip_IP_IPv6_1 through client_ip_IP_IPv6_8 |
8 | Expanded IPv6 groups converted from hexadecimal to decimal, each between 0 and 65535 |
client_ip_IP_Missing |
1 | 1 for missing, otherwise 0 |
The table below shows the exact numeric values constructed for the three prediction rows after omitting constant training features. It is transposed for readability: each row names a separate model input column, and every value is a scalar number.
| Source input column | Model input column | IPv4 row: 192.0.2.200 |
IPv6 row: 2001:db8::c8 |
Missing IP row |
|---|---|---|---|---|
client_ip |
client_ip_IP_Version |
4 | 6 | 0 |
client_ip |
client_ip_IP_IPv4_1 |
192 | 0 | 0 |
client_ip |
client_ip_IP_IPv4_3 |
2 | 0 | 0 |
client_ip |
client_ip_IP_IPv4_4 |
200 | 0 | 0 |
client_ip |
client_ip_IP_IPv6_1 |
0 | 8193 | 0 |
client_ip |
client_ip_IP_IPv6_2 |
0 | 3512 | 0 |
client_ip |
client_ip_IP_IPv6_8 |
0 | 200 | 0 |
client_ip |
client_ip_IP_Missing |
0 | 0 | 1 |
request_count |
request_count (unchanged) |
12 | 7 | 3 |
client_ip_IP_IPv4_2 and client_ip_IP_IPv6_3 through
client_ip_IP_IPv6_7 are omitted because they are zero for every training row.
The original client_ip string column is removed from model inputs. No feature
contains a list, and request_count remains a single numeric column. Its values
are independent of whether the IP address is IPv4, IPv6, or missing.
For IPv6, :: expands to the omitted zero groups. The hexadecimal groups 2001,
db8, and c8 become decimal values 8193, 3512, and 200 respectively.
Unused components for the other address version are zero-filled. The missing
indicator, when retained, distinguishes missing addresses from valid all-zero addresses such as
0.0.0.0 and ::.
Constant columns are dropped after transformation¶
After expanding an IP column, AutoML keeps only generated columns with at least two distinct values in the learner's training data. Constant columns are dropped before scaling. This includes version, individual components, and the missing indicator; none of them is always kept.
For example, if training contains only 192.0.2.1, 192.0.2.2, and 192.0.2.3:
| Candidate feature | Training values | Outcome |
|---|---|---|
client_ip_IP_Version |
Always 4 |
Dropped |
client_ip_IP_IPv4_1 |
Always 192 |
Dropped |
client_ip_IP_IPv4_2 |
Always 0 |
Dropped |
client_ip_IP_IPv4_3 |
Always 2 |
Dropped |
client_ip_IP_IPv4_4 |
1, 2, 3 |
Kept |
| All eight IPv6 components | Always 0 |
Dropped |
client_ip_IP_Missing |
Always 0 |
Dropped |
Each learner saves its selection from its training fold. Validation, prediction, and reload reuse exactly those columns. Dropped columns are never added back, even if new data would make them vary. Retained columns are not dropped when a prediction batch happens to contain constant values.
A new IPv6 or missing address is accepted, but only the retained features reach the model. In this IPv4-only example, no version, IPv6, or missing-indicator feature is available to distinguish those cases.
The Decision Tree in this example uses these values directly. Algorithms that
require scaling use scaling learned from the training data. You continue to pass
raw addresses to predict(); AutoML constructs the numeric columns internally.
Run the script¶
From a repository checkout with mljar-supervised installed:
python examples/scripts/binary_classifier_ip_addresses.py --results-path AutoML_IP
The complete script checks prediction equality after reload. Use a fresh results directory for a new training run.
Supported inputs¶
- Every non-missing training value must be a valid plain IPv4 or IPv6 string.
- Surrounding whitespace is ignored; equivalent IPv6 spellings are accepted.
- Explicit pandas
categorycolumns retain categorical handling. To use that behavior, assignX["client_ip"] = X["client_ip"].astype("category")before fit. - Mixed IP/non-IP training columns follow existing categorical/text handling.
- Once detected as IP, malformed prediction values raise a column/row error.
- CIDR notation, ports, and IPv6 zone identifiers are excluded.
See Preprocessing for generated features and missing-value behavior, and Save and Load models for model persistence.