load: reach a VL model's layers, not just load it - #104
Merged
Merged
Conversation
Fixing the loader moved the failure two lines down. pollard-probe loaded Qwen2-VL successfully and then died on `len(model.model.layers)` -- which reads as a different bug and is the same one: a vision-language model keeps its decoder under model.language_model, so every direct `model.model.layers` raises AttributeError on a model that has just loaded fine. Seven tools did this: probe, abliterate, doctor, lowbit, gptq, hf_smooth, palette. All go through pollard_load.text_layers() now, which finds the stack for either family. VL: 28 layers, first block Qwen2VLDecoderLayer, attn projections reachable One of those was a latent crash unrelated to VL: pollard_abliterate imported the helper inside main(), while abliterate() uses it at module scope -- so the tool worked when driven from the CLI and raised NameError when called as a library. Hoisted. A test walks every pollard_*.py and fails on a direct layer reach, the same way the loader test does, because the two halves of this bug are easy to fix separately and leave half-broken. 58/58 pass.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixing the loader moved the failure two lines down. pollard-probe loaded Qwen2-VL successfully and then died on
len(model.model.layers)-- which reads as a different bug and is the same one: a vision-language model keeps its decoder under model.language_model, so every directmodel.model.layersraises AttributeError on a model that has just loaded fine.Seven tools did this: probe, abliterate, doctor, lowbit, gptq, hf_smooth, palette. All go through pollard_load.text_layers() now, which finds the stack for either family.
VL: 28 layers, first block Qwen2VLDecoderLayer, attn projections reachable
One of those was a latent crash unrelated to VL: pollard_abliterate imported the helper inside main(), while abliterate() uses it at module scope -- so the tool worked when driven from the CLI and raised NameError when called as a library. Hoisted.
A test walks every pollard_*.py and fails on a direct layer reach, the same way the loader test does, because the two halves of this bug are easy to fix separately and leave half-broken.
58/58 pass.