MCPcopy Create free account
hub / github.com/QWTforGithub/T2LDM / build_model

Function build_model

models/CLIP/clip/model.py:455–492  ·  view source on GitHub ↗
(state_dict: dict)

Source from the content-addressed store, hash-verified

453
454
455def build_model(state_dict: dict):
456 vit = "visual.proj" in state_dict
457
458 if vit:
459 vision_width = state_dict["visual.conv1.weight"].shape[0]
460 vision_layers = len([k for k in state_dict.keys() if k.startswith("visual.") and k.endswith(".attn.in_proj_weight")])
461 vision_patch_size = state_dict["visual.conv1.weight"].shape[-1]
462 grid_size = round((state_dict["visual.positional_embedding"].shape[0] - 1) ** 0.5)
463 image_resolution = vision_patch_size * grid_size
464 else:
465 counts: list = [len(set(k.split(".")[2] for k in state_dict if k.startswith(f"visual.layer{b}"))) for b in [1, 2, 3, 4]]
466 vision_layers = tuple(counts)
467 vision_width = state_dict["visual.layer1.0.conv1.weight"].shape[0]
468 output_width = round((state_dict["visual.attnpool.positional_embedding"].shape[0] - 1) ** 0.5)
469 vision_patch_size = None
470 assert output_width ** 2 + 1 == state_dict["visual.attnpool.positional_embedding"].shape[0]
471 image_resolution = output_width * 32
472
473 embed_dim = state_dict["text_projection"].shape[1]
474 context_length = state_dict["positional_embedding"].shape[0]
475 vocab_size = state_dict["token_embedding.weight"].shape[0]
476 transformer_width = state_dict["ln_final.weight"].shape[0]
477 transformer_heads = transformer_width // 64
478 transformer_layers = len(set(k.split(".")[2] for k in state_dict if k.startswith("transformer.resblocks")))
479
480 model = CLIP(
481 embed_dim,
482 image_resolution, vision_layers, vision_width, vision_patch_size,
483 context_length, vocab_size, transformer_width, transformer_heads, transformer_layers
484 )
485
486 for key in ["input_resolution", "context_length", "vocab_size"]:
487 if key in state_dict:
488 del state_dict[key]
489
490 convert_weights(model)
491 model.load_state_dict(state_dict)
492 return model.eval()

Callers 1

loadFunction · 0.70

Calls 3

CLIPClass · 0.85
convert_weightsFunction · 0.85
load_state_dictMethod · 0.45

Tested by

no test coverage detected