MCPcopy Create free account

hub / github.com/adamkarvonen/activation_oracles / functions

Functions716 in github.com/adamkarvonen/activation_oracles

↓ 1 callersFunction_find_all_grade_paths_by_kind_and_mi
( agent_root: Path, organism: str, model: str, *, mi: int, is_baseline: bool, llm_
experiments/final_paper_plots/plot_em_agent_multi_bars.py:105
↓ 1 callersFunction_find_all_grade_paths_by_kind_and_mi
( agent_root: Path, organism: str, model: str, *, mi: int, is_baseline: bool, llm_
experiments/final_paper_plots/plot_em_agent.py:97
↓ 1 callersFunction_gather_model_stats
(model: str, highlight_keyword: str)
experiments/final_paper_plots/plot_personaqa_sequence_vs_token.py:43
↓ 1 callersFunction_gender_stats
()
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:77
↓ 1 callersFunction_gender_stats
()
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:83
↓ 1 callersMethod_get_qa_for_sentence
(self, sentence, sentence_entities, all_entities, num_qa_per_sample)
nl_probes/dataset_classes/classification_dataset_manager.py:199
↓ 1 callersFunction_inner
()
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:705
↓ 1 callersFunction_legend_labels
Convert LoRA names to human-readable labels.
experiments/final_paper_plots/plot_all_data_diversity.py:426
↓ 1 callersFunction_legend_labels
Convert model names to human-readable labels.
experiments/final_paper_plots/plot_personaqa_knowledge_eval_all_models.py:145
↓ 1 callersFunction_legend_labels
Convert LoRA names to human-readable labels.
experiments/final_paper_plots/plot_personaqa_results_all_models.py:418
↓ 1 callersFunction_load_export_records
(export_dir: str)
experiments/final_paper_plots/plot_agent_adam.py:301
↓ 1 callersFunction_load_export_records
(export_dir: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:301
↓ 1 callersFunction_load_export_records
(export_dir: str)
experiments/final_paper_plots/plot_em_agent.py:289
↓ 1 callersFunction_model_display_name
(model: str)
experiments/final_paper_plots/plot_agent_adam.py:60
↓ 1 callersFunction_model_display_name
(model: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:60
↓ 1 callersFunction_model_display_name
(model: str)
experiments/final_paper_plots/plot_em_agent.py:52
↓ 1 callersFunction_name_matches
(name: str)
experiments/final_paper_plots/plot_agent_adam.py:122
↓ 1 callersFunction_name_matches
(name: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:122
↓ 1 callersFunction_name_matches
(name: str)
experiments/final_paper_plots/plot_em_agent.py:114
↓ 1 callersFunction_normalize_behavior_item
Return a 7-tuple corresponding to the unified behavior format. Order matches the original repo logic: (system, control_user, control_thought,
nl_probes/dataset_classes/misc/latentqa_loader.py:95
↓ 1 callersFunction_normalize_llm_id
(llm_id: str)
experiments/final_paper_plots/plot_agent_adam.py:74
↓ 1 callersFunction_normalize_llm_id
(llm_id: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:74
↓ 1 callersFunction_normalize_llm_id
(llm_id: str)
experiments/final_paper_plots/plot_em_agent.py:66
↓ 1 callersFunction_personaqa_open_ended_stats
Get PersonaQA open-ended stats for a single model.
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:131
↓ 1 callersFunction_plot_results_panel
Plot a single panel with bars using shared palette.
experiments/final_paper_plots/plot_all_data_diversity.py:502
↓ 1 callersFunction_plot_results_panel
Plot a single panel with bars using shared palette.
experiments/final_paper_plots/plot_personaqa_knowledge_eval_all_models.py:165
↓ 1 callersFunction_plot_results_panel
Plot a single panel with bars using shared palette.
experiments/final_paper_plots/plot_personaqa_results_all_models.py:524
↓ 1 callersFunction_print_preview
(prompts_by_group: dict[str, list[dict[str, str]]], per_group_preview: int = 2)
experiments/explorations/zero_shot_classification_eval.py:63
↓ 1 callersFunction_qwen_taboo_stats
()
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:65
↓ 1 callersFunction_qwen_taboo_stats
()
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:71
↓ 1 callersFunction_results_root_for_agent_type
(cfg, agent_type: str)
experiments/final_paper_plots/plot_agent_adam.py:66
↓ 1 callersFunction_results_root_for_agent_type
(cfg, agent_type: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:66
↓ 1 callersFunction_results_root_for_agent_type
(cfg, agent_type: str)
experiments/final_paper_plots/plot_em_agent.py:58
↓ 1 callersFunction_ssc_stats
Load SSC results, filtering to only the highlight keyword file.
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:89
↓ 1 callersFunction_ssc_stats
Load SSC results, filtering to only the highlight keyword file.
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:95
↓ 1 callersFunction_strip
(obj)
nl_probes/dataset_classes/act_dataset_manager.py:38
↓ 1 callersFunction_style_highlight
(bar, color=INTERP_BAR_COLOR, hatch="////")
experiments/final_paper_plots/plot_qwen3-8b_eval_results.py:234
↓ 1 callersFunction_style_highlight
Style the highlighted bar with hatch and edge.
experiments/final_paper_plots/plot_all_data_diversity.py:439
↓ 1 callersFunction_style_highlight
Style the highlighted bar with edge (no hatch).
experiments/final_paper_plots/plot_personaqa_knowledge_eval_all_models.py:158
↓ 1 callersFunction_style_highlight
Style the highlighted bar with hatch and edge.
experiments/final_paper_plots/plot_classification_eval_all_models.py:265
↓ 1 callersFunction_style_highlight
Style the highlighted bar with hatch and edge.
experiments/final_paper_plots/plot_personaqa_results_all_models.py:431
↓ 1 callersFunction_taboo_stats
()
experiments/final_paper_plots/plot_secret_keeping_sequence_vs_token.py:55
↓ 1 callersFunction_taboo_stats
()
experiments/final_paper_plots/plot_combined_sequence_vs_token.py:61
↓ 1 callersFunction_variant_params
(v_key: str)
experiments/final_paper_plots/plot_agent_adam.py:212
↓ 1 callersFunction_variant_params
(v_key: str)
experiments/final_paper_plots/plot_em_agent_multi_bars.py:212
↓ 1 callersFunction_variant_params
(v_key: str)
experiments/final_paper_plots/plot_em_agent.py:204
↓ 1 callersFunctionanalyse_quirk
( records: list[Record], response_type: ResponseType = "token_responses", best_of_n: int = 5, )
experiments/final_paper_plots/plot_ssc_results.py:224
↓ 1 callersFunctionanalyse_quirk
( records: list[Record], response_type: ResponseType = "token_responses", best_of_n: int = 5 )
experiments/final_paper_plots/plot_layer_comparison_secret_keeping.py:298
↓ 1 callersFunctionanalyse_quirk
( records: list[Record], response_type: ResponseType = "token_responses", best_of_n: int = 5 )
experiments/final_paper_plots/plot_secret_keeping_results.py:348
↓ 1 callersFunctionanalyse_quirk
( records: list[Record], response_type: ResponseType = "token_responses", best_of_n: int = 5, )
experiments/plotting/plot_ssc_results.py:176
↓ 1 callersFunctionanalyse_quirk
( records: list[Record], response_type: ResponseType = "token_responses", best_of_n: int = 5 )
experiments/plotting/plot_secret_keeping_results.py:311
↓ 1 callersMethodas_text
(self)
nl_probes/autointerp_detection_eval/caller.py:35
↓ 1 callersFunctionasync_judge
(judge_fn)
datasets/latentqa_datasets/curate_gpt_data.py:109
↓ 1 callersFunctionasync_query
( model_name, user_prompt, system_prompt="", max_tokens=100, temperature=1.0, stop_seq
datasets/latentqa_datasets/curate_gpt_data.py:34
↓ 1 callersFunctionasync_vote
(query, args)
datasets/latentqa_datasets/curate_gpt_data.py:104
↓ 1 callersFunctionbuild_datasets
( cfg: SelfInterpTrainingConfig, dataset_loaders: list[ActDatasetLoader], max_len_percentile: floa
nl_probes/sft.py:533
↓ 1 callersFunctionbuild_detection_prompts
Zips explanations with their SAETrainTest records and creates prompts for the detector. Includes an assertion that feature indices stay align
nl_probes/autointerp_detection_eval/lora_hf_eval.py:163
↓ 1 callersFunctionbuild_eval_data
Builds the generator's evaluation dataset given a single SAE reference and feature ids. Fixed the original bug by passing device and dtype ex
nl_probes/autointerp_detection_eval/lora_hf_eval.py:85
↓ 1 callersFunctionbuild_loader_groups
( *, model_name: str, layer_percents: list[int], act_collection_batch_size: int, save_acts
nl_probes/sft.py:600
↓ 1 callersFunctionbuild_loader_groups_no_model
Build dataset loaders without model_kwargs (no model needed for loading existing data).
experiments/explorations/count_training_data.py:48
↓ 1 callersFunctionbuild_prompts_and_targets
Create chat prompts and targets for a single dataset file.
experiments/patchscopes/patchscopes_full_open_ended_eval.py:136
↓ 1 callersFunctionbuild_prompts_and_targets
Create chat prompts and targets for a single dataset file.
experiments/patchscopes/patchscopes_zero_shot_all.py:64
↓ 1 callersFunctionbuild_zero_shot_prompts
Build a mapping from dataset group -> list of zero-shot prompts, using the same CLASSIFICATION_DATASETS config structure as experiments/class
experiments/explorations/zero_shot_classification_eval.py:15
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/patchscopes/plot_patchscopes_results.py:78
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/final_paper_plots/plot_gender_eval_results.py:92
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/final_paper_plots/plot_old_personaqa_results.py:79
↓ 1 callersFunctioncalculate_accuracy
(record: dict, investigator_lora: str | None)
experiments/final_paper_plots/plot_taboo_eval_results.py:90
↓ 1 callersFunctioncalculate_accuracy
Calculate accuracy for specified datasets.
experiments/final_paper_plots/plot_all_data_diversity.py:115
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/final_paper_plots/plot_personaqa_results.py:94
↓ 1 callersFunctioncalculate_accuracy
Calculate accuracy for a record. Args: record: The record containing responses offset: Token offset for token-level accuracy
experiments/final_paper_plots/plot_personaqa_open_ended_results.py:117
↓ 1 callersFunctioncalculate_accuracy
Calculate accuracy for a record using model-specific offset. Args: record: The record containing responses offset: Token offset f
experiments/final_paper_plots/plot_personaqa_results_all_models.py:188
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/plotting/plot_gender_eval_results.py:80
↓ 1 callersFunctioncalculate_accuracy
(record: dict, investigator_lora: str)
experiments/plotting/plot_taboo_eval_results.py:75
↓ 1 callersFunctioncalculate_accuracy
(record)
experiments/plotting/plot_personaqa_results.py:63
↓ 1 callersFunctioncalculate_confidence_interval
Calculate 95% confidence interval for accuracy data.
experiments/final_paper_plots/plot_personaqa_results_all_models.py:404
↓ 1 callersFunctioncalculate_confidence_interval_binomial
Calculate binomial confidence interval for accuracy.
experiments/final_paper_plots/plot_all_data_diversity.py:131
↓ 1 callersFunctioncalculate_confidence_interval_binomial
Calculate binomial confidence interval for accuracy based on sample count.
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:734
↓ 1 callersFunctioncalculate_personaqa_accuracy
Calculate accuracy for PersonAQA record using sequence-based open-ended matching.
experiments/final_paper_plots/plot_all_data_diversity.py:248
↓ 1 callersFunctioncalculate_taboo_accuracy
Calculate accuracy for Taboo record - exactly matching original script logic.
experiments/final_paper_plots/plot_all_data_diversity.py:322
↓ 1 callersFunctioncall_model_for_sae_explanation
Call the specified model to get an explanation for the SAE feature.
nl_probes/autointerp_detection_eval/eval_detection_v2.py:379
↓ 1 callersFunctioncanonical_dataset_id
Strip 'classification_' prefix if present so keys match your IID/OOD lists.
experiments/classification_eval.py:133
↓ 1 callersFunctioncheck_answer_match
Check if the answer matches the ground truth, handling ambiguous cases.
experiments/personaqa_knowledge_eval.py:84
↓ 1 callersFunctioncheck_answer_match
Check if the answer matches the ground truth, handling ambiguous cases (for open-ended).
experiments/final_paper_plots/plot_all_data_diversity.py:232
↓ 1 callersFunctioncheck_answer_match
Check if the answer matches the ground truth, handling ambiguous cases (for open-ended).
experiments/final_paper_plots/plot_model_progression_line_chart_shapes.py:183
↓ 1 callersFunctioncheck_personaqa_match
Check if the answer matches the ground truth.
experiments/final_paper_plots/plot_lr_sweep_combined.py:169
↓ 1 callersFunctionci95
(values: list[float])
experiments/final_paper_plots/plot_layer_comparison_secret_keeping.py:381
↓ 1 callersFunctionci95
(values: list[float])
experiments/final_paper_plots/plot_qwen3-8b_eval_results.py:184
↓ 1 callersFunctionci95
(values: list[float])
experiments/plotting/plot_secret_keeping_results.py:424
↓ 1 callersFunctioncollect_activations_with_lora
Collect activations with LoRA enabled, disabled, and the difference.
experiments/simple_llama_model_demo.py:162
↓ 1 callersFunctioncollect_activations_without_lora
( model: AutoModelForCausalLM, submodules: dict, inputs_BL: dict[str, torch.Tensor], act_layer
experiments/simple_llama_model_demo.py:142
↓ 1 callersFunctioncollect_activations_without_lora
( model: AutoModelForCausalLM, submodules: dict, inputs_BL: dict[str, torch.Tensor], )
experiments/patchscopes/patchscopes_full_open_ended_eval.py:223
↓ 1 callersFunctioncollect_past_lens_acts
( dataset_config: DatasetLoaderConfig, custom_dataset_params: PastLensDatasetConfig, tokenizer: Au
nl_probes/dataset_classes/past_lens_dataset.py:169
↓ 1 callersFunctioncollect_target_activations
( model: AutoModelForCausalLM, inputs_BL: dict[str, torch.Tensor], config: VerbalizerEvalConfig,
nl_probes/base_experiment.py:322
↓ 1 callersFunctioncollect_target_responses
( model: AutoModelForCausalLM, tokenizer: AutoTokenizer, context_prompts: list[list[dict[str, str]
nl_probes/base_experiment.py:288
↓ 1 callersFunctioncombine_with_ultrachat
Sample from UltraChat, filter to first turn only, filter by max character length from the taboo dataset, then combine and shuffle with the ma
nl_probes/trl_training/taboo_train.py:308
↓ 1 callersFunctioncompare_patchscope_responses
(response: str, target_response: str)
experiments/patchscopes/simple_patchscopes_demo.py:247
↓ 1 callersFunctioncompare_patchscope_responses
(response: str, target_response: str)
experiments/patchscopes/patchscopes_zero_shot_all.py:58
↓ 1 callersFunctioncompare_patchscope_responses
(response: str, target_response: str)
experiments/patchscopes/patchscopes_zero_shot.py:175
↓ 1 callersFunctioncompute_sae_activations_for_sentences
Compute SAE activations for a list of sentences and return SentenceInfo objects.
nl_probes/autointerp_detection_eval/create_hard_negatives_v2.py:76
↓ 1 callersFunctionconstruct_train_dataset
( custom_dataset_params: SAEExplanationDatasetConfig, dataset_type: str, dataset_size: int, la
nl_probes/dataset_classes/sae_training_data.py:556
← previousnext →201–300 of 716, ranked by callers