HTB: Spin Glass Brain Challenge

Spin Glass Brain - HackTheBox Challenge Writeup

Challenge Information

FieldValue
NameSpin Glass Brain
CategoryMisc (AI/ML)
DifficultyHard
Authord3vn0mi

Description

You are provided with the network weights of a Hopfield Network with discrete bipolar units (i.e. each neuron can have a value -1 or 1). Can you obtain the patterns that the model has been trained with? Note: the patterns are encoded with indices 1-16.

Solution

The challenge ships a Jupyter notebook plus a weight matrix for a classic (Hopfield-style) associative memory network. Recovering the “training patterns” means finding the stable attractor states of the network and reading off which one corresponds to each requested index — the resulting patterns turn out to be images (brain MRI slices), each stamped with one character of the flag.

The interesting part of this challenge isn’t the neural network math — it’s that a naive “clamp the index bits and settle” approach silently fails, because the index encoding occupies a vanishingly small fraction of the network’s neurons and therefore exerts almost no pull on the dynamics. The real trick is a two-phase partial-cue recall: clamp on the index bits to steer the network into the right basin, then release the clamp and let it settle freely so the network’s own dynamics confirm (and clean up) the recovered pattern.

Key Steps

1. Recover the real challenge artifact.

The staged challenge.ipynb was 0 bytes — a known packaging quirk. The original password-protected archive next to it held the real files.

Terminal window
# staged notebook was empty; pull the real bundle from the original zip
unzip -o -P hackthebox /out/<uuid>.zip
# yields:
# challenge.ipynb (2,197 bytes)
# weights.npy (327 MB)

2. Inspect the weight matrix.

import numpy as np
W = np.load('weights.npy', mmap_mode='r')
print(W.shape, W.dtype) # (6400, 6400) float64
print(np.allclose(W, W.T)) # symmetric -> classic Hopfield weight matrix
print(np.diag(W).max()) # 0 -> zero self-connections

A 6400×6400 symmetric zero-diagonal matrix is exactly the weight structure of a Hopfield network storing 80×80-pixel bipolar (±1) images as attractor states.

3. Read the notebook’s encoding scheme.

The notebook’s encode_pattern(key, arr) function overwrites the first 6 neurons of each stored pattern with the 6-bit binary encoding of that pattern’s index (0–15 fits in 6 bits with room to spare). So every stored memory carries its own index label baked into a tiny corner of the 6400-neuron state vector.

4. First attempt (naive clamp) — and why it failed.

import numpy as np
def clamp_and_settle(W, index_bits, n_neurons=6400, clamp_size=6, iters=200):
rng = np.random.default_rng(0)
x = rng.choice([-1, 1], size=n_neurons) # random start
x[:clamp_size] = index_bits # clamp index neurons
for _ in range(iters):
x = np.sign(W @ x)
x[:clamp_size] = index_bits # re-clamp every step
return x

Clamping just 6 out of 6400 neurons exerts a negligible field on the overall energy landscape — a random start plus a hard clamp on 0.1% of the state just falls into whichever basin the random initialization was already closest to. This produced duplicate/garbage patterns with no relation to the requested index.

5. The working technique — two-phase partial-cue recall.

import numpy as np
N, CLAMP = 6400, 6
def recall(W, index_bits, iters=200):
# Phase 1: cue purely from the index-bit sub-block of W, clamped
x = np.sign(W[:, :CLAMP] @ index_bits)
for _ in range(iters):
x = np.sign(W @ x)
x[:CLAMP] = index_bits # keep steering toward the right basin
# Phase 2: release the clamp, run FREE dynamics to let the network
# settle into the true, self-consistent attractor.
for _ in range(iters):
x = np.sign(W @ x)
# Self-consistency check: the network should have restored the
# original index bits on its own if we landed in the right basin.
recovered_index = x[:CLAMP]
return x, recovered_index

The key insight: using the full column W[:, :CLAMP] to construct the initial cue (instead of a random vector) gives the network a real signal correlated with the target memory, rather than 6 clamped bits fighting against 6394 random ones. Releasing the clamp afterward lets genuine Hopfield energy-minimization dynamics pull the state the rest of the way into the nearest stored attractor — and if the index bits it independently reconstructs match the one requested, that’s confirmation the correct pattern was recovered.

6. Sweep the remaining stragglers.

This got 14 of 16 patterns directly. The last few needed noisy restarts of the cue vector and a batched random-restart sweep (many parallel random seeds feeding into the same recall procedure, keeping any run whose self-reported index matched) to escape local minima / spurious basins.

import numpy as np
def sweep(W, target_index_bits, n_seeds=400, iters=200):
rng = np.random.default_rng()
for seed in range(n_seeds):
noise = rng.choice([-1, 1], size=N) * 0.0 # placeholder for jitter
x, recovered = recall(W, target_index_bits + noise)
if np.array_equal(recovered, target_index_bits):
return x
return None

7. Render and read the flag.

import numpy as np, matplotlib.pyplot as plt
patterns = {} # index -> 6400-length ±1 vector
for idx in range(16):
patterns[idx] = recall_for_index(idx) # via steps 5-6
fig, axes = plt.subplots(4, 4, figsize=(10, 10))
for idx, ax in zip(range(16), axes.flat):
img = (patterns[idx].reshape(80, 80) + 1) / 2 # -1/1 -> 0/1 for display
ax.imshow(img, cmap='gray')
ax.set_title(str(idx))
plt.savefig('flag_patterns.png')

Each recovered 80×80 pattern was a brain MRI slice with a single character overlaid. Reading them off in index order (0 through 15 — the description says indices 1–16, but index 0 actually holds the first character, H) spelled out the flag.

Tools Used

  • Python 3 / NumPy — Hopfield network dynamics (sign, matrix-vector products), bipolar state vectors
  • SciPy (scipy.sparse.linalg.eigsh) — attempted spectral analysis of the weight matrix during exploration
  • Matplotlib — rendering recovered ±1 vectors as 80×80 grayscale images to read the embedded characters
  • unzip — recovering the real notebook/weights from the password-protected original archive (hackthebox)
  • Background shell tasks for long-running random-restart sweeps, polled asynchronously

Key Learnings

  • A Hopfield network’s storage capacity is diffuse across the whole state vector. Clamping a tiny fraction of neurons (6 of 6400 here) does almost nothing to bias which attractor a random start converges to — the field it exerts is far too weak relative to the network’s dominant eigenmodes.
  • Partial-cue recall works far better as a two-phase process: first construct a cue that’s actually correlated with the target memory (e.g., derived from the relevant weight sub-block, W[:, :k] @ cue_bits, rather than pure noise), steer with it, and only then release the clamp for free energy-minimization dynamics to finish the job.
  • Self-labeling patterns make a great correctness oracle. Because each stored pattern encoded its own index in a reserved sub-block of neurons, recall attempts could self-verify: if the network’s freely-settled state didn’t reproduce the requested index bits, the recall had converged to the wrong basin (or a spurious/mixture state) and needed a retry.
  • Symmetric, zero-diagonal weight matrices are a strong fingerprint for classic (Hebbian-rule) Hopfield networks — checking W == W.T and diag(W) == 0 early confirms the architecture before investing in dynamics-based recovery.
  • When staged CTF artifacts show up empty or truncated, check for an original password-protected archive alongside them before assuming the challenge itself is broken — a common hackthebox-password zip pattern recurs across these challenges.
  • Random-restart batching pays off for stubborn attractors. A handful of patterns resisted the primary recall method entirely; running many parallel random-seeded attempts and filtering by the self-consistency check (recovered index bits == requested index) reliably picked off the stragglers without needing a smarter single-shot algorithm.