MCPcopy Create free account

hub / github.com/adamkarvonen/activation_oracles / functions

Functions716 in github.com/adamkarvonen/activation_oracles

↓ 1 callersFunctionload_results
Load all JSON files from the directory.
experiments/final_paper_plots/plot_personaqa_results.py:117
↓ 1 callersFunctionload_results
Load all JSON files from the directory. Args: json_dir: Directory containing JSON files offset: Token offset for token-level accu
experiments/final_paper_plots/plot_personaqa_open_ended_results.py:152
↓ 1 callersFunctionload_results
Load all JSON files from the directory.
experiments/plotting/plot_gender_eval_results.py:102
↓ 1 callersFunctionload_results
Load all JSON files from the directory.
experiments/plotting/plot_ssc_results.py:289
↓ 1 callersFunctionload_results
Load all JSON files from the directory.
experiments/plotting/plot_taboo_eval_results.py:103
↓ 1 callersFunctionload_results
Load all JSON files from the directory.
experiments/plotting/plot_personaqa_results.py:87
↓ 1 callersFunctionload_results_from_folder
Load all JSON results from folder and calculate accuracies keyed by LoRA name.
experiments/final_paper_plots/plot_classification_eval.py:140
↓ 1 callersFunctionload_results_from_folder
Load all JSON results from folder and calculate accuracies keyed by LoRA name.
experiments/final_paper_plots/plot_classification_eval_all_models.py:179
↓ 1 callersFunctionload_results_from_folder
Load all JSON results from folder and calculate accuracies keyed by LoRA name.
experiments/plotting/plot_classification_eval.py:125
↓ 1 callersFunctionload_sae_data_from_sft_data_file
( dataset_config: DatasetLoaderConfig, custom_dataset_params: SAEExplanationDatasetConfig, tokeniz
nl_probes/dataset_classes/sae_training_data.py:604
↓ 1 callersFunctionload_split_acc
(folder: Path, keywords)
experiments/final_paper_plots/plot_classification_layer_sweep_lines.py:64
↓ 1 callersFunctionload_ssc_results
Load all JSON files from the directory.
experiments/plotting/plot_secret_keeping_results.py:391
↓ 1 callersFunctionload_ssc_results_sync
(json_dir: str)
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:704
↓ 1 callersFunctionload_taboo_gemma_results
Load Taboo results for Gemma-2-9B-IT, matching the secret-keeping plot logic (taboo_calculate_accuracy with CHOSEN_TABOO_PROMPT and TABOO_SEQ
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:366
↓ 1 callersFunctionload_taboo_qwen_results
Load Taboo results for Qwen3-8B (open-ended, one-word secret). Mirrors plot_taboo_eval_results.py for Qwen3-8B with SEQUENCE=False and the
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:265
↓ 1 callersFunctionload_taboo_results
Load all JSON files from the directory.
experiments/final_paper_plots/plot_qwen3-8b_eval_results.py:145
↓ 1 callersFunctionload_taboo_results
Load Taboo results from directory - exactly matching original script logic.
experiments/final_paper_plots/plot_all_data_diversity.py:350
↓ 1 callersFunctionload_taboo_results
Load all JSON files from the directory.
experiments/plotting/plot_secret_keeping_results.py:121
↓ 1 callersFunctionload_training_data_only
Load only training data from existing dataset files.
experiments/explorations/count_training_data.py:172
↓ 1 callersFunctionmain
()
experiments/patchscopes/patchscopes_zero_shot_all.py:125
↓ 1 callersFunctionmain
()
experiments/patchscopes/plot_patchscopes_results.py:311
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:203
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_gender_eval_results.py:421
↓ 1 callersFunctionmain
Main function to process all JSON files.
experiments/final_paper_plots/fix_taboo_json_files.py:62
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_old_personaqa_results.py:313
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_train_loss_lr_sweep.py:70
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_ssc_results.py:662
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_classification_layer_sweep_lines.py:188
↓ 1 callersFunctionmain
(tasks: list[str])
experiments/final_paper_plots/plot_layer_comparison_secret_keeping.py:433
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_qwen3-8b_eval_results.py:302
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_personaqa_sequence_vs_token.py:130
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_taboo_eval_results.py:427
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_all_data_diversity.py:623
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_lr_sweep_combined.py:274
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_secret_keeping_results.py:603
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_classification_eval.py:310
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_personaqa_results.py:398
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_personaqa_open_ended_results.py:439
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_personaqa_knowledge_eval_all_models.py:303
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_classification_eval_all_models.py:627
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:883
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_personaqa_results_all_models.py:706
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_llama_layer0_vs_layer1_latentqa.py:220
↓ 1 callersFunctionmain
()
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:224
↓ 1 callersFunctionmain
()
experiments/plotting/plot_gender_eval_results.py:341
↓ 1 callersFunctionmain
()
experiments/plotting/plot_ssc_results.py:517
↓ 1 callersFunctionmain
()
experiments/plotting/plot_taboo_eval_results.py:379
↓ 1 callersFunctionmain
()
experiments/plotting/plot_secret_keeping_results.py:553
↓ 1 callersFunctionmain
()
experiments/plotting/plot_classification_eval.py:291
↓ 1 callersFunctionmain
()
experiments/plotting/plot_personaqa_results.py:268
↓ 1 callersFunctionmain
()
experiments/explorations/count_training_data.py:229
↓ 1 callersFunctionmain
Main function to process SAE activations, get explanations, and run precision/recall evaluation. Args: sae_file: Path to the SAE har
nl_probes/autointerp_detection_eval/eval_detection_v2.py:940
↓ 1 callersFunctionmain
( target_features: Sequence[int] = [0, 1, 2, 3, 4, 5, 6, 7, 8, 9], target_sentences: int = 20, top
nl_probes/autointerp_detection_eval/create_hard_negatives_v2.py:130
↓ 1 callersFunctionmain
()
nl_probes/autointerp_detection_eval/lora_hf_eval.py:302
↓ 1 callersFunctionmake_debug_collator
Wrap a collator to print tokenization details for debugging.
nl_probes/trl_training/personaqa_train.py:26
↓ 1 callersFunctionmake_random_explanation
(items: Slist[SAETrainTestWithExplanation], name: str)
nl_probes/autointerp_detection_eval/eval_detection_v2.py:927
↓ 1 callersFunctionmanual_qwen3_assistant_mask
Create a mask where 1 indicates assistant tokens and 0 indicates non-assistant tokens. Args: tokenized: Dictionary containing 'input
nl_probes/trl_training/taboo_train.py:144
↓ 1 callersFunctionoom_preflight_check
( cfg: SelfInterpTrainingConfig, training_data: list[TrainingDataPoint], model: AutoModelForCausal
nl_probes/sft.py:278
↓ 1 callersFunctionparse_args
()
datasets/latentqa_datasets/curate_gpt_data.py:146
↓ 1 callersFunctionparse_yes_no_qas
Parse an LLM response containing four Q/A pairs in this form: <question> ... </question> <answer> ... </answer> Returns a list
nl_probes/dataset_classes/sae_training_data.py:357
↓ 1 callersFunctionpersonaqa_calculate_accuracy
(record, sequence: bool)
experiments/final_paper_plots/plot_qwen3-8b_eval_results.py:69
↓ 1 callersFunctionplot_all_eval_types
Create a single plot with all three eval types as subplots.
experiments/final_paper_plots/plot_all_data_diversity.py:544
↓ 1 callersFunctionplot_all_models
()
experiments/final_paper_plots/plot_classification_layer_sweep_lines.py:115
↓ 1 callersFunctionplot_all_models
Create plot with all three models as subplots. Args: all_results: List of result dictionaries for each model model_names: List of
experiments/final_paper_plots/plot_personaqa_knowledge_eval_all_models.py:202
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/final_paper_plots/plot_gender_eval_results.py:281
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/final_paper_plots/plot_ssc_results.py:522
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/final_paper_plots/plot_taboo_eval_results.py:287
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/plotting/plot_gender_eval_results.py:235
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/plotting/plot_ssc_results.py:437
↓ 1 callersFunctionplot_by_keyword_with_extras
Plot exactly one LoRA (selected by required_keyword in its name) plus extra bars. Asserts that exactly one LoRA matches and that extra_bars h
experiments/plotting/plot_taboo_eval_results.py:252
↓ 1 callersFunctionplot_comparison
(tasks: list[tuple[str, float, float, float, float]], out_path: Path)
experiments/final_paper_plots/plot_llama_layer0_vs_layer1_latentqa.py:166
↓ 1 callersFunctionplot_iid_and_ood
Create separate IID and OOD plots with highlighted LoRA.
experiments/final_paper_plots/plot_classification_eval.py:301
↓ 1 callersFunctionplot_iid_and_ood
Create separate IID and OOD plots with highlighted LoRA.
experiments/plotting/plot_classification_eval.py:278
↓ 1 callersFunctionplot_per_word_accuracy
Create separate plots for each investigator showing per-word accuracy.
experiments/patchscopes/plot_patchscopes_results.py:256
↓ 1 callersFunctionplot_per_word_accuracy
Create separate plots for each investigator showing per-word accuracy.
experiments/final_paper_plots/plot_gender_eval_results.py:371
↓ 1 callersFunctionplot_per_word_accuracy
Create separate plots for each investigator showing per-word accuracy.
experiments/final_paper_plots/plot_ssc_results.py:612
↓ 1 callersFunctionplot_per_word_accuracy
Create separate plots for each investigator showing per-word accuracy.
experiments/plotting/plot_gender_eval_results.py:299
↓ 1 callersFunctionplot_progression_lines
Plot one line per series, with: - X-axis = MODEL_TYPE_ORDER - Y-axis = accuracy / quirk score - Color determined by evaluation
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:764
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA.
experiments/patchscopes/plot_patchscopes_results.py:158
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/final_paper_plots/plot_gender_eval_results.py:172
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/final_paper_plots/plot_old_personaqa_results.py:161
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/final_paper_plots/plot_ssc_results.py:420
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/final_paper_plots/plot_taboo_eval_results.py:185
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/plotting/plot_gender_eval_results.py:157
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/plotting/plot_ssc_results.py:346
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA, highlighting exactly one LoRA.
experiments/plotting/plot_taboo_eval_results.py:161
↓ 1 callersFunctionplot_results
Create a bar chart of average accuracy by investigator LoRA.
experiments/plotting/plot_personaqa_results.py:147
↓ 1 callersFunctionplot_sequence_vs_token
(stats: list[tuple[str, float, float, float, float]])
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:141
↓ 1 callersFunctionplot_sequence_vs_token
(model_data: list[tuple[str, list[tuple[str, float, float, float, float]]]])
experiments/final_paper_plots/plot_personaqa_sequence_vs_token.py:64
↓ 1 callersFunctionplot_sequence_vs_token
(stats: list[tuple[str, float, float, float, float]])
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:163
↓ 1 callersFunctionprint_trainable_parameters
(model)
nl_probes/trl_training/personaqa_train.py:59
↓ 1 callersFunctionprint_trainable_parameters
(model)
nl_probes/trl_training/taboo_train.py:34
↓ 1 callersFunctionproportion_confidence
Compute proportion statistics. Returns (p, se, lower, upper) - p: proportion correct (in [0,1]) - se: standard error of the proporti
nl_probes/utils/eval.py:186
↓ 1 callersFunctionread_jsonl_file_into_basemodel
( path: Path | str, basemodel: type[GenericBaseModel], limit: int | None = None )
nl_probes/autointerp_detection_eval/caller.py:246
↓ 1 callersFunctionread_jsonl_file_into_basemodel_async
( path: AnyioPath, basemodel: type[GenericBaseModel] )
nl_probes/autointerp_detection_eval/caller.py:383
↓ 1 callersFunctionread_sae_file
Read SAEs from a JSONL file with optional start index and limit. Args: sae_file: Path to JSONL file. limit: Maximum number of SAE
nl_probes/autointerp_detection_eval/eval_detection_v2.py:56
↓ 1 callersFunctionreorder_by_labels
Reorder bars: highlight first, then Data Matched, then LatentQA + Classification, then others.
experiments/final_paper_plots/plot_all_data_diversity.py:471
↓ 1 callersFunctionreorder_by_labels
Reorder bars: Full Dataset -> SPQA + Classification -> SPQA Only (Pan et al.) -> Classification -> Original Model.
experiments/final_paper_plots/plot_personaqa_results_all_models.py:491
↓ 1 callersMethodreplace_explanation
(self, explanation: ChatHistory, explainer_model: str)
nl_probes/autointerp_detection_eval/eval_detection_v2.py:331
↓ 1 callersFunctionreplace_model_name
(name)
datasets/latentqa_datasets/curate_gpt_data.py:162
← previousnext →401–500 of 716, ranked by callers