HTB: AI SPACE Challenge
AI SPACE - HackTheBox Challenge Writeup
Challenge Information
| Field | Value |
|---|---|
| Name | AI SPACE |
| Category | Misc (AI/ML) |
| Difficulty | Easy |
| Author | d3vn0mi |
Description
You are assigned the important mission of locating and identifying the infamous space hacker. Your investigation begins by analyzing the data patterns and breach points identified in the latest cyber-attacks. Use the provided coordinates of the last known signal origins to narrow down his potential hideouts. Utilize advanced tracking algorithms to follow the digital footprint left by the hacker.
The challenge ships a single artifact, distance_matrix.npy — a NumPy array described only as “coordinates of the last known signal origins.” No source code, no service to talk to; the entire challenge is a data-forensics puzzle.
Solution
The provided archive extracted to a 0-byte distance_matrix.npy despite the zip listing a healthy ~26 MB entry for it — a sign the extract had silently failed rather than the file genuinely being empty. The archive turned out to be password-protected with the classic HTB default password, hackthebox, and re-extracting with that password recovered the real file.
Loading the recovered array showed it was (1808, 1808), float64, symmetric, with a zero diagonal — the textbook shape of a pairwise distance matrix over 1,808 points, matching the brief’s talk of “coordinates of last known signal origins.”
To go from distances between points back to actual point positions, the standard technique is classical multidimensional scaling (MDS): double-center the squared distance matrix and take its eigendecomposition. The eigenvalues came out as roughly 8225.4, 41.6, ~1e-12, ... — a massive drop-off after the second value, meaning the point cloud is exactly two-dimensional. That’s a strong signal an image (or text) had been encoded into the distances.
Plotting the top two eigenvector-derived components as a scatter plot rendered the flag text directly — literal ASCII-art typography made of 1,808 dots. The very first render looked like mirrored gibberish; MDS output has a per-axis sign ambiguity (eigenvectors are only defined up to sign), so negating the second component un-mirrored the text into readable characters. Zooming into thirds of the plot confirmed the leetspeak glyph choices (d1st4nt uses a 4 for the “a”, but spac3 keeps a plain “a”).
Key Steps
1. Recover the real artifact from the password-protected zip:
# Initial extract silently produced a 0-byte distance_matrix.npy# even though the zip's file listing showed ~26MB for that entry.# Re-extract with the standard HTB default password:unzip -o -P hackthebox /out/<challenge>.zip2. Fingerprint the recovered array:
import numpy as np
D = np.load('distance_matrix.npy')print(D.shape) # (1808, 1808)print(D.dtype) # float64print(np.allclose(D, D.T)) # True -> symmetricprint(np.allclose(np.diag(D), 0)) # True -> zero diagonal# Conclusion: this is a pairwise distance matrix over 1808 points.3. Invert distances to coordinates with classical MDS:
import numpy as np
D = np.load('distance_matrix.npy')n = D.shape[0]D2 = D ** 2
# Double-centering matrixJ = np.eye(n) - np.ones((n, n)) / n
# Gram matrix via double-centering: B = -1/2 * J * D^2 * JB = -0.5 * J @ D2 @ J
# Eigendecompositioneigvals, eigvecs = np.linalg.eigh(B)
# Sort descending — top eigenvalues reveal the intrinsic dimensionalityorder = np.argsort(eigvals)[::-1]eigvals = eigvals[order]eigvecs = eigvecs[:, order]
print(eigvals[:5])# ~[8225.4, 41.6, 1e-12, ...] -> effectively rank-2:# the point cloud lives in a flat 2D plane.
# Reconstruct 2D coordinates from the top-2 componentsX = eigvecs[:, :2] * np.sqrt(eigvals[:2])np.save('/tmp/X.npy', X)4. Plot the reconstructed points to reveal the flag:
import numpy as npimport matplotlibmatplotlib.use('Agg')import matplotlib.pyplot as plt
X = np.load('/tmp/X.npy')
# First render was mirrored/gibberish -- MDS has a per-axis sign# ambiguity (eigenvectors only defined up to sign). Flip the# second axis to un-mirror the text into readable characters.X[:, 1] *= -1
fig, ax = plt.subplots(figsize=(20, 4))ax.scatter(X[:, 0], X[:, 1], s=2, c='black')ax.set_aspect('equal')ax.axis('off')plt.savefig('flag.png', dpi=200, bbox_inches='tight')The resulting scatter plot read out the flag as literal typography rendered in point clouds.
Tools Used
- Python 3 / NumPy — array loading, matrix operations, eigendecomposition
- matplotlib — rendering the reconstructed point cloud
unzip— recovering the password-protected artifact- Classical Multidimensional Scaling (MDS) — the core technique for inverting a distance matrix back into 2D point coordinates
Key Learnings
- A 0-byte file after extraction doesn’t mean an empty file — check the archive listing first. If the zip’s own metadata shows a nontrivial size for an entry that extracted empty, suspect a password gate or extraction failure before assuming the artifact is broken.
- Always try the HTB default password (
hackthebox) on protected archives before reaching for OSINT or brute-forcing — many challenges gate artifacts this way as a trivial speed bump rather than a real puzzle. - A pairwise distance matrix can be inverted to coordinates via classical MDS (double-center the squared distances, eigendecompose, keep the top-k eigenvectors scaled by
sqrt(eigenvalue)). This is a general technique worth recognizing: symmetric, zero-diagonal matrices are a strong “this is a distance matrix” fingerprint. - Check the eigenvalue spectrum before plotting — a sharp drop-off after the first few eigenvalues reveals the true intrinsic dimensionality of the encoded data (here, exactly 2D), which is the signal that a flat image/text had been steganographically encoded into pairwise distances.
- MDS reconstructions have sign ambiguity per axis — if a decoded plot looks like mirrored noise, try flipping individual axes before assuming the technique failed.