Quick Start Guide¶
setselector requires a runtime CSV and a feature CSV, or one evaluation XML
file produced by the potassco-benchmark-tool.
Additionally, a runtime cutoff and the desired number of instances are required.
Select from CSV files¶
Small example files can be found here.
They are also included in the repository in example/.
To select using CSV files, run:
The command prints the selected instance names to standard output. With
--log=info, it also reports the available data, filtering, runtime summaries,
sample features, and cluster statistics.
Input format¶
The runtime CSV has one instance per row. Its first column is the instance name and the remaining columns are solver runtimes. The header identifies the solvers:
The feature CSV uses the same instance names. Its first column is the instance name and the remaining columns are numeric feature values. The header identifies the features:
Rows are joined by instance name. Rows without matching data, with inconsistent
feature lengths, or with NaN or infinite feature values are ignored. Runtime
values above the cutoff are treated as timeouts at the cutoff. Negative feature
values are limited to -1, which represents a missing feature value.
Select from evaluation XML¶
Alternatively, an evaluation XML file can be used instead of the two CSV files:
By default, the features rules_s and bodies_s are extracted. Other
measures can be selected with a comma-separated list:
Repeated runs of an instance are combined using their median feature values and median runtime. Missing evaluation runs are treated as timeouts.
Input format¶
Evaluation files must be provided in the format generated by the
btool eval command from the potassco-benchmark-tool.
Filtering and sampling¶
Before sampling, the tool removes instances for which every solver reaches the
cutoff and instances solved by every solver below easyK * cutoff. This behavior
can be overridden using --keep-too-hard or --keep-too-easy.
The default sampler uses a Gaussian runtime distribution, averages runtimes across solvers, and allows each cluster to represent at most 20% of the requested sample. The relevant options are:
| Option | Default | Description |
|---|---|---|
--aggregate |
avg |
Use avg, min, or ind runtime hardness. |
--dist |
gauss |
Use gauss, uni, exp, or log sampling. |
--frac |
0.2 |
Maximum fraction of the sample from one cluster. |
--reps |
100 |
Number of k-means repetitions. |
--easyK |
0.1 |
Easy-instance threshold as a fraction of the cutoff. |
Use --split to randomly remove half the instances as a test set before
sampling. The selection algorithm uses a fixed seed, so the same input and
options produce the same selection.
Info
The returned number of instances can be smaller than --n when the available
data or cluster limits do not permit the requested sample size.