Skip to content

Quick Start Guide

setselector requires a runtime CSV and a feature CSV, or one evaluation XML file produced by the potassco-benchmark-tool. Additionally, a runtime cutoff and the desired number of instances are required.

Select from CSV files

Small example files can be found here. They are also included in the repository in example/. To select using CSV files, run:

setselector --runtimes ./example/times.csv --features ./example/features.csv --n 5 --cutoff 300

The command prints the selected instance names to standard output. With --log=info, it also reports the available data, filtering, runtime summaries, sample features, and cluster statistics.

Input format

The runtime CSV has one instance per row. Its first column is the instance name and the remaining columns are solver runtimes. The header identifies the solvers:

instances,s1,s2
inst01,5,8
inst02,12,15

The feature CSV uses the same instance names. Its first column is the instance name and the remaining columns are numeric feature values. The header identifies the features:

instances,f1,f2
inst01,3,12
inst02,5,10

Rows are joined by instance name. Rows without matching data, with inconsistent feature lengths, or with NaN or infinite feature values are ignored. Runtime values above the cutoff are treated as timeouts at the cutoff. Negative feature values are limited to -1, which represents a missing feature value.

Select from evaluation XML

Alternatively, an evaluation XML file can be used instead of the two CSV files:

setselector --eval eval.xml --n 5 --cutoff 300

By default, the features rules_s and bodies_s are extracted. Other measures can be selected with a comma-separated list:

setselector --eval eval.xml --eval-features atoms,basic_rules --n 5 --cutoff 300

Repeated runs of an instance are combined using their median feature values and median runtime. Missing evaluation runs are treated as timeouts.

Input format

Evaluation files must be provided in the format generated by the btool eval command from the potassco-benchmark-tool.

Filtering and sampling

Before sampling, the tool removes instances for which every solver reaches the cutoff and instances solved by every solver below easyK * cutoff. This behavior can be overridden using --keep-too-hard or --keep-too-easy.

The default sampler uses a Gaussian runtime distribution, averages runtimes across solvers, and allows each cluster to represent at most 20% of the requested sample. The relevant options are:

Option Default Description
--aggregate avg Use avg, min, or ind runtime hardness.
--dist gauss Use gauss, uni, exp, or log sampling.
--frac 0.2 Maximum fraction of the sample from one cluster.
--reps 100 Number of k-means repetitions.
--easyK 0.1 Easy-instance threshold as a fraction of the cutoff.

Use --split to randomly remove half the instances as a test set before sampling. The selection algorithm uses a fixed seed, so the same input and options produce the same selection.

Info

The returned number of instances can be smaller than --n when the available data or cluster limits do not permit the requested sample size.