MCPcopy Create free account
hub / github.com/Open-Quant/openquant / plot_feature_importance

Function plot_feature_importance

crates/openquant/src/feature_importance.rs:191–205  ·  view source on GitHub ↗
(
    importance: &BTreeMap<String, ImportanceStats>,
    oob_score: f64,
    oos_score: f64,
    output_path: Option<&str>,
)

Source from the content-addressed store, hash-verified

189/// is fitted on the training rows (with their `sample_weight`), the test rows are scored, and
190/// then each feature column is shuffled in turn within the test rows and scored again. The
191/// per-fold importance is `(base - perm) / (0 - perm)` for [`Scoring::NegLogLoss`] and
192/// `(base - perm) / (1 - perm)` for [`Scoring::Accuracy`] and [`Scoring::F1`]; it is 0 when
193/// the denominator is 0 or the ratio is not finite. Test-fold scores **are** weighted by
194/// `sample_weight`. Accuracy and F1 use [`SimpleClassifier::predict`] (so an override is
195/// honoured), negative log loss uses [`SimpleClassifier::predict_proba`]. The result is the mean
196/// over folds and its standard error (sample deviation, ddof 1, over `sqrt(n_folds)`).
197///
198/// 1 means damaging the feature destroyed everything the model had, 0 that the model did not
199/// need it, negative that it did better without it. Because the shuffle stays within each test
200/// fold, a very persistent feature is somewhat understated.
201///
202/// # Errors
203///
204/// - [`FeatureImportanceError::EmptyXy`] if `x` or `y` is empty.
205/// - [`FeatureImportanceError::XyLengthMismatch`] if `x.len() != y.len()`.
206/// - [`FeatureImportanceError::LengthMismatch`] (`"feature_names"`) if the first row of `x`
207/// does not have one entry per feature name.
208/// - [`FeatureImportanceError::RaggedX`] if the rows of `x` differ in length.

Calls

no outgoing calls

Tested by 1