HTB: The Art of Capture Challenge
The Art of Capture - HackTheBox Challenge Writeup
Challenge Information
| Field | Value |
|---|---|
| Name | The Art of Capture |
| Category | Forensics |
| Difficulty | Hard |
| Author | d3vn0mi |
Description
I had a dream that my PC was captured by evil daemons. The next day, strange things started happening to it. There are two parts of the flag in this challenge.
The challenge ships a packet capture of a compromised host talking to a C2 server, plus supporting artifacts. Two independent halves of the flag are buried in two different pieces of evidence — a decrypted screenshot and a decrypted executable — and both need to be extracted and reconstructed before the flag can be assembled.
Solution
Overview
The pcap contains a JSON-based C2 protocol running on 192.168.1.34:4444, exposing /api/v2/login, /api/v2/ping, and /api/v2/query endpoints. The login response leaks an obfuscated AES key; the real session key is recoverable from a memory dump (DbgInfo.DMP) included with the challenge. Once the key is known, the query request/response bodies (base64 + AES-128-CTR + gzip) can be decrypted, revealing:
- A desktop screenshot (PNG) containing the first half of the flag rendered as figlet ASCII art.
- A gzip’d PE (
rev.exe) which RC4-decrypts a shellcode stager at runtime, containing the second half of the flag in cleartext.
Each half required a different extraction technique — deterministic pixel-grid OCR for the ASCII-art screenshot, and static/dynamic RC4 decryption for the binary.
Key Steps
1. Recover the AES session key from the C2 traffic and memory dump
The login response in the pcap returns an obfuscated key blob tied to a session id:
{"id": "gxksd3v7", "k": "<obfuscated-key-blob>"}Cross-referencing this with strings pulled from the memory dump recovered the real AES-128 session key used for all subsequent traffic:
Session key: kVUboRSaneIPXWg1Task and result bodies on /api/v2/query are base64(AES-128-CTR(gzip(payload))). Decrypting the captured query responses with this key and inflating the gzip stream yields the screenshot PNG and the PE payload.
2. Extract the screenshot and confirm it hides ASCII-art text
from PIL import Image
im = Image.open('screenshot.png')print(im.size) # full desktop screenshot
# Crop to the region containing the terminal window with the flagim.convert('RGB').crop((0, 60, 1280, 210)).resize( (1280 * 2, 150 * 2), Image.LANCZOS).save('p1a.png')Zooming in showed a terminal rendering large figlet-style ASCII-art text — too small/aliased to read reliably by eye. Manual eyeballing produced ambiguous reads (e.g. confusing 1S for 1$, W4t for Wat), which is exactly the kind of error that needed to be eliminated deterministically.
3. Binarize the screenshot onto its terminal cell grid
Rather than trust human OCR, the ASCII-art region was reduced to its underlying character-cell grid and then binarized to isolate the actual “pixels” the figlet font renderer had drawn:
from PIL import Imageimport numpy as np
a = np.array(Image.open('screenshot.png').convert('L'))dark = (a < 128) # thresholded / binarized rendering
# Locate the bounding rows/cols of the ASCII-art blockrows = np.where(dark.any(axis=1))[0]cols = np.where(dark[83:193].any(axis=0))[0]Autocorrelation on the row/column pixel-density profile recovered the terminal’s per-character cell size (5 px wide × 10 px tall). Snapping the bitmap to that grid produced a clean, fixed-size matrix of cells:
d = (a < 128).astype(np.uint8)grid = []for r in range(11): # 11 rows of figlet output row_cells = [] for c in range(256): # up to 256 columns cell = d[r*10:(r+1)*10, x0+c*5 : x0+(c+1)*5] row_cells.append(cell.tobytes()) grid.append(row_cells)Deduplicating the cell bitmaps found there were only 6 distinct glyphs in use: space, $, /, \, |, _ — consistent with a figlet “block” style font. Each cell was mapped to its corresponding character and reassembled into a raw art.txt file, an exact character-for-character transcript of the rendered ASCII art (no more guesswork).
4. Brute-force match art.txt against figlet fonts to recover the plaintext
With a perfect (not human-read) transcript of the ASCII art in hand, the next step was to reverse figlet’s rendering: find which font and which source string reproduces art.txt exactly.
import pyfiglet
# Font fingerprinting: try candidate fonts until block shapes matchfor font in ['bigmoney-ne', 'big_money-ne', 'block', 'banner3', ...]: f = pyfiglet.Figlet(font=font) # compare rendered glyph shapes against known cells from art.txtbig_money-ne was identified as the matching font. From there, a greedy character-by-character search reconstructed the source string:
import pyfiglet, string
H = 11F = pyfiglet.Figlet(font='big_money-ne', width=100000) # see gotcha belowart = [l.rstrip() for l in open('art.txt').read().split('\n')][:H]
charset = string.ascii_letters + string.digits + '{}_$!@#'known = "HTB{REDACTED}while True: for ch in charset: candidate = known + ch rendered = [l.rstrip() for l in F.renderText(candidate).split('\n')] if matches_prefix(rendered, art): known = candidate break else: break # no character extends the match furtherGotcha:
pyfiglet.figlet_format()defaults to a render width of 80 columns and silently wraps long output into stacked blocks below a certain length, corrupting any naive column-by-column comparison. This must be overridden explicitly withFiglet(font=..., width=100000)so the full string renders on one unwrapped line.
Greedily matching the full 11×256 character grid against big_money-ne output recovered flag part 1 exactly, with no OCR ambiguity:
HTB{REDACTED} part 1 — redacted>_5. Unpack rev.exe and recover the RC4-encrypted stager
The second decrypted C2 artifact was a gzip’d Windows PE (rev.exe). Static analysis showed the PE decrypts an embedded blob with RC4 at runtime and copies the plaintext into RWX memory before executing it — a classic self-decrypting shellcode loader pattern.
strings -n 4 stage2.bin # inspect decrypted stage for embedded strings/configxxd stage2.bin | head -50 # confirm shellcode prologue / stager signatureThe decrypted stage resolved to a Metasploit-style reverse_winhttp stager (WinINet-based reverse shell shellcode) configured to callback to 192.168.56.1. Critically, the flag’s second half was appended as cleartext directly after the shellcode body in the decrypted buffer — no further decoding needed:
d = open('stage2.bin', 'rb').read()i = d.index(b'St4y') # locate flag-tail marker in decrypted stagepart2 = d[i:].decode()print(part2)# -> "<flag part 2 — redacted>}"6. Assemble and submit the full flag
part1 = "HTB{REDACTED} part 1>"part2 = "<flag part 2>}"flag = part1 + part2print(flag)HTB{REDACTED}Submitted via the HTB API for challenge 766 (The Art of Capture, Forensics):
from htb_api import HTBClientc = HTBClient()c.submit_challenge(766, "HTB{REDACTED}", difficulty=4)# -> {"message": "Congratulations!"}Tools Used
- Wireshark / pcap analysis — extracting the JSON C2 protocol (
/api/v2/login|ping|query) - Python
cryptography— AES-128-CTR decryption of C2 payloads gzip— inflating decrypted C2 payload bodies- Memory forensics on
DbgInfo.DMP— recovering the real AES session key Pillow/numpy— pixel-level image analysis, grid detection via autocorrelation, and binarization of the screenshotpyfiglet— brute-force font/string matching to reverse the ASCII-art renderingstrings/xxd— static inspection of the decrypted PE stage- RC4 decryption (manual/Python) — unpacking the self-decrypting PE loader
- Metasploit stager knowledge — recognizing the
reverse_winhttpshellcode pattern
Key Learnings
- Never trust human OCR on figlet/ASCII-art renders. Small aliasing artifacts (
1Svs1$,W4tvsWat) are easy to misread by eye. Converting the image to a strict character-cell grid, binarizing each cell, and deduplicating glyph bitmaps turns “read the picture” into an exact, verifiable transcription problem. - Reversing figlet output is a font+width matching problem, not a text-recognition problem. Once you have an exact glyph transcript, brute-forcing candidate fonts and then greedily matching characters against
Figlet(...).renderText()output reconstructs the original string with certainty — provided the render width is set high enough to avoid line-wrapping artifacts (width=100000vs the default 80). - Self-decrypting PE loaders (RC4-into-RWX) are a common obfuscation pattern for embedding raw shellcode/stagers inside a benign-looking executable; dumping the decrypted buffer at runtime (or replicating the RC4 routine statically) exposes both the payload’s C2 configuration and any appended plaintext (here, the flag).
- Multi-part flags spread across unrelated artifact types (image vs. binary) require correlating multiple decryption pipelines (C2 session key → AES-CTR → gzip for the screenshot; PE unpacking → RC4 for the executable) before either half becomes visible — treat each artifact’s encoding as a separate mini-challenge within the larger one.