HTB: The Art of Capture Challenge

The Art of Capture - HackTheBox Challenge Writeup

Challenge Information

FieldValue
NameThe Art of Capture
CategoryForensics
DifficultyHard
Authord3vn0mi

Description

I had a dream that my PC was captured by evil daemons. The next day, strange things started happening to it. There are two parts of the flag in this challenge.

The challenge ships a packet capture of a compromised host talking to a C2 server, plus supporting artifacts. Two independent halves of the flag are buried in two different pieces of evidence — a decrypted screenshot and a decrypted executable — and both need to be extracted and reconstructed before the flag can be assembled.

Solution

Overview

The pcap contains a JSON-based C2 protocol running on 192.168.1.34:4444, exposing /api/v2/login, /api/v2/ping, and /api/v2/query endpoints. The login response leaks an obfuscated AES key; the real session key is recoverable from a memory dump (DbgInfo.DMP) included with the challenge. Once the key is known, the query request/response bodies (base64 + AES-128-CTR + gzip) can be decrypted, revealing:

  • A desktop screenshot (PNG) containing the first half of the flag rendered as figlet ASCII art.
  • A gzip’d PE (rev.exe) which RC4-decrypts a shellcode stager at runtime, containing the second half of the flag in cleartext.

Each half required a different extraction technique — deterministic pixel-grid OCR for the ASCII-art screenshot, and static/dynamic RC4 decryption for the binary.

Key Steps

1. Recover the AES session key from the C2 traffic and memory dump

The login response in the pcap returns an obfuscated key blob tied to a session id:

{"id": "gxksd3v7", "k": "<obfuscated-key-blob>"}

Cross-referencing this with strings pulled from the memory dump recovered the real AES-128 session key used for all subsequent traffic:

Session key: kVUboRSaneIPXWg1

Task and result bodies on /api/v2/query are base64(AES-128-CTR(gzip(payload))). Decrypting the captured query responses with this key and inflating the gzip stream yields the screenshot PNG and the PE payload.

2. Extract the screenshot and confirm it hides ASCII-art text

from PIL import Image
im = Image.open('screenshot.png')
print(im.size) # full desktop screenshot
# Crop to the region containing the terminal window with the flag
im.convert('RGB').crop((0, 60, 1280, 210)).resize(
(1280 * 2, 150 * 2), Image.LANCZOS
).save('p1a.png')

Zooming in showed a terminal rendering large figlet-style ASCII-art text — too small/aliased to read reliably by eye. Manual eyeballing produced ambiguous reads (e.g. confusing 1S for 1$, W4t for Wat), which is exactly the kind of error that needed to be eliminated deterministically.

3. Binarize the screenshot onto its terminal cell grid

Rather than trust human OCR, the ASCII-art region was reduced to its underlying character-cell grid and then binarized to isolate the actual “pixels” the figlet font renderer had drawn:

from PIL import Image
import numpy as np
a = np.array(Image.open('screenshot.png').convert('L'))
dark = (a < 128) # thresholded / binarized rendering
# Locate the bounding rows/cols of the ASCII-art block
rows = np.where(dark.any(axis=1))[0]
cols = np.where(dark[83:193].any(axis=0))[0]

Autocorrelation on the row/column pixel-density profile recovered the terminal’s per-character cell size (5 px wide × 10 px tall). Snapping the bitmap to that grid produced a clean, fixed-size matrix of cells:

d = (a < 128).astype(np.uint8)
grid = []
for r in range(11): # 11 rows of figlet output
row_cells = []
for c in range(256): # up to 256 columns
cell = d[r*10:(r+1)*10, x0+c*5 : x0+(c+1)*5]
row_cells.append(cell.tobytes())
grid.append(row_cells)

Deduplicating the cell bitmaps found there were only 6 distinct glyphs in use: space, $, /, \, |, _ — consistent with a figlet “block” style font. Each cell was mapped to its corresponding character and reassembled into a raw art.txt file, an exact character-for-character transcript of the rendered ASCII art (no more guesswork).

4. Brute-force match art.txt against figlet fonts to recover the plaintext

With a perfect (not human-read) transcript of the ASCII art in hand, the next step was to reverse figlet’s rendering: find which font and which source string reproduces art.txt exactly.

import pyfiglet
# Font fingerprinting: try candidate fonts until block shapes match
for font in ['bigmoney-ne', 'big_money-ne', 'block', 'banner3', ...]:
f = pyfiglet.Figlet(font=font)
# compare rendered glyph shapes against known cells from art.txt

big_money-ne was identified as the matching font. From there, a greedy character-by-character search reconstructed the source string:

import pyfiglet, string
H = 11
F = pyfiglet.Figlet(font='big_money-ne', width=100000) # see gotcha below
art = [l.rstrip() for l in open('art.txt').read().split('\n')][:H]
charset = string.ascii_letters + string.digits + '{}_$!@#'
known = "HTB{REDACTED}
while True:
for ch in charset:
candidate = known + ch
rendered = [l.rstrip() for l in F.renderText(candidate).split('\n')]
if matches_prefix(rendered, art):
known = candidate
break
else:
break # no character extends the match further

Gotcha: pyfiglet.figlet_format() defaults to a render width of 80 columns and silently wraps long output into stacked blocks below a certain length, corrupting any naive column-by-column comparison. This must be overridden explicitly with Figlet(font=..., width=100000) so the full string renders on one unwrapped line.

Greedily matching the full 11×256 character grid against big_money-ne output recovered flag part 1 exactly, with no OCR ambiguity:

HTB{REDACTED} part 1 — redacted>_

5. Unpack rev.exe and recover the RC4-encrypted stager

The second decrypted C2 artifact was a gzip’d Windows PE (rev.exe). Static analysis showed the PE decrypts an embedded blob with RC4 at runtime and copies the plaintext into RWX memory before executing it — a classic self-decrypting shellcode loader pattern.

strings -n 4 stage2.bin # inspect decrypted stage for embedded strings/config
xxd stage2.bin | head -50 # confirm shellcode prologue / stager signature

The decrypted stage resolved to a Metasploit-style reverse_winhttp stager (WinINet-based reverse shell shellcode) configured to callback to 192.168.56.1. Critically, the flag’s second half was appended as cleartext directly after the shellcode body in the decrypted buffer — no further decoding needed:

d = open('stage2.bin', 'rb').read()
i = d.index(b'St4y') # locate flag-tail marker in decrypted stage
part2 = d[i:].decode()
print(part2)
# -> "<flag part 2 — redacted>}"

6. Assemble and submit the full flag

part1 = "HTB{REDACTED} part 1>"
part2 = "<flag part 2>}"
flag = part1 + part2
print(flag)
HTB{REDACTED}

Submitted via the HTB API for challenge 766 (The Art of Capture, Forensics):

from htb_api import HTBClient
c = HTBClient()
c.submit_challenge(766, "HTB{REDACTED}", difficulty=4)
# -> {"message": "Congratulations!"}

Tools Used

  • Wireshark / pcap analysis — extracting the JSON C2 protocol (/api/v2/login|ping|query)
  • Python cryptography — AES-128-CTR decryption of C2 payloads
  • gzip — inflating decrypted C2 payload bodies
  • Memory forensics on DbgInfo.DMP — recovering the real AES session key
  • Pillow / numpy — pixel-level image analysis, grid detection via autocorrelation, and binarization of the screenshot
  • pyfiglet — brute-force font/string matching to reverse the ASCII-art rendering
  • strings / xxd — static inspection of the decrypted PE stage
  • RC4 decryption (manual/Python) — unpacking the self-decrypting PE loader
  • Metasploit stager knowledge — recognizing the reverse_winhttp shellcode pattern

Key Learnings

  • Never trust human OCR on figlet/ASCII-art renders. Small aliasing artifacts (1S vs 1$, W4t vs Wat) are easy to misread by eye. Converting the image to a strict character-cell grid, binarizing each cell, and deduplicating glyph bitmaps turns “read the picture” into an exact, verifiable transcription problem.
  • Reversing figlet output is a font+width matching problem, not a text-recognition problem. Once you have an exact glyph transcript, brute-forcing candidate fonts and then greedily matching characters against Figlet(...).renderText() output reconstructs the original string with certainty — provided the render width is set high enough to avoid line-wrapping artifacts (width=100000 vs the default 80).
  • Self-decrypting PE loaders (RC4-into-RWX) are a common obfuscation pattern for embedding raw shellcode/stagers inside a benign-looking executable; dumping the decrypted buffer at runtime (or replicating the RC4 routine statically) exposes both the payload’s C2 configuration and any appended plaintext (here, the flag).
  • Multi-part flags spread across unrelated artifact types (image vs. binary) require correlating multiple decryption pipelines (C2 session key → AES-CTR → gzip for the screenshot; PE unpacking → RC4 for the executable) before either half becomes visible — treat each artifact’s encoding as a separate mini-challenge within the larger one.