COBOLSTACK.COM
//MCP     EXEC PGM=AGENT,PARM='VERIFIED.NOT.AUTONOMOUS'

MCP servers for the estate: an agent that can run, not just read

Two Model Context Protocol servers — one over a z/OS-compatible mainframe, one over an IBM i companion — give your coding agents and LLM tooling a place where COBOL compiles, JCL runs through a JES2-class spool, CICS-style transactions answer, VSAM clusters hold records, and every byte comes back with its EBCDIC and packed-decimal truth intact. The docs are public; access is issued per organisation.

//LOOP    DD DSN=AGENT.WORK.LOOP,DISP=SHR

The loop an agent actually gets

// one modernization-validation pass, tool by tool (SteelFrame, z/OS side)
create_workspace      → private copy of the estate; the live one never changes
upload_source         → COBOL + copybooks + JCL into PDS members
check_source          → inline compile check; diagnostics back as data
compile_program       → compile + link through a real JES2-class job; listing returned
run_program           → DDs declared as JSON; SYSOUT + return code back
run_batch             → whole job: joblog, every SYSOUT, every output dataset, clock pinned
cics_run_script       → a 3270 transcript, one AID key per step, fields by name
decode_records        → records through the copybook, beside the raw EBCDIC/COMP-3 bytes
export_state          → datasets, DB2-class tables, spool: hex + decoded rows + checksums
// then compare against YOUR expected results. The server never judges
// equivalence and never converts code — it returns what the machine holds and does.
MCP2 SERVERSREADY
//DOCS    DD DSN=TWO.SERVERS.TWO.DOCS

Two servers, two documents

Each server speaks MCP over stdio (a small bridge) or plain HTTP (POST /mcp, JSON-RPC 2.0, bearer token). Both run the same tool surface whether you point them at the persistent estate or at a private, RAM-only workspace. Summaries below; the linked pages are the authority.

ServerWhat the tools coverDocs
SteelFrame
mainframe · z/OS-compatible
Source & compile: check_source, upload_source, compile_program, run_program, one-shot compile-and-run for the private lane. JES2: submit_job, wait_job, read_spool, hold/release/cancel/purge, WTOR replies. Data: datasets and members, define_cluster, vsam_info, read_records/write_records, decode_records/encode_records, sql (SPUFI-style), GDG staging. Analysis: search_source, find_program, explain_program, inventory_scan, analysis_graph/impact/metrics, bms_preview. Online: headless CICS-class sessions, cics_run_script, cics_link with COMMAREA, CSD read/define. Oracle runs: run_batch, run_capture, export_state, export_bundle; CA-7-style and Control-M-style scheduler cycles. Isolation: create_workspace / reset_workspace / destroy_workspace. Mainframe (z/OS) MCP docs ↗
SteelFrame X
IBM i · AS/400-compatible
Find & describe: libraries, objects, members, read_source, describe_file, describe_program, ddl_generate, analyze. Seed & build: upload_source, compile (CRTBNDRPG, CRTBNDCBL, CRTCLPGM, CRTPF, CRTDSPF…), load_file, SAVF upload/download. Run: call_program with parameter write-back, run_command (the CL command set), submit_job, job_log, spool_read, sql for DB2 for i. Screens: 5250 sessions one AID key at a time, fields by name; screen_run_script. Batch & schedule: run_batch, run_chain on a single-initiator queue, job schedule entries, job queues. Records: decode_records through the DDS layout beside both byte images, export_state. Isolation: workspaces, library lists, authority ceilings on every key (*READ, *RUN, *ALL). IBM i (AS/400) MCP docs ↗
The one rule both servers keep The server returns what the machine holds and does. It never judges equivalence and never converts code. Your harness — or ours, under engagement — does the comparing. That keeps the evidence trail honest and the agent out of the judge's chair.
//EVAL    DD DSN=WHAT.AN.EVALUATION.LOOKS.LIKE

What an evaluation looks like

An enterprise logistics group evaluating the platform this quarter set the scope in one sentence: analyse IBM COBOL applications and related mainframe assets, and validate a modernization workflow with synthetic samples — proving COBOL compile/link/run, JCL/JES2 job execution and spool, CICS-style transactions, VSAM organizations, and EBCDIC/COMP-3 behaviour against expected results. That is the shape most evaluations take, so here is each item mapped to what the servers actually do.

See it run by an outside AI. We handed this exact list to Grok, xAI's desktop app, over MCP and let it write its own tests: 45 passes, 1 skip and 1 fail, whose cause was on our side and has since been fixed. Screenshots and results per area: We Let Grok Test SteelFrame Over MCP →
You want to validateThe agent callsWhat comes back to compare
COBOL compile / link / run upload_source → check_source → compile_program → run_program compiler diagnostics as data; the compile-and-link job's listing and return code; the run's SYSOUT, return code and every output dataset it wrote. Copybooks, subprograms, intrinsic functions and byte-true COMP-3/zoned storage are part of the pipeline, not a mode.
JCL / JES2 execution + spool submit_job → wait_job → read_spool; or run_batch in one call JESMSGLG, JESJCL, JESYSMSG and every SYSOUT DD; condition codes per step; DISP, COND and IF/THEN/ELSE, PROCs with symbolics and overrides, referbacks, GDGs. run_batch adds the output datasets and can pin the system clock so date-driven logic reproduces exactly.
CICS-style transactions cics_open_session → cics_send → cics_screen; cics_run_script; cics_link a 3270 screen after each AID key with fields addressed by BMS name; a full transcript for a scripted flow; the COMMAREA after a LINK; CSD definitions read back. Pseudo-conversational patterns, TS/TD queues and channels/containers behave per the published semantics.
VSAM organizations define_cluster, IDCAMS via JCL, vsam_info, read_records / write_records KSDS, ESDS, RRDS and LDS with alternate indexes and paths; cluster attributes and statistics; records by key, by RBA, by relative number — decoded through the copybook on request. Your program's file status codes are the ones the spec says.
EBCDIC / COMP-3 behaviour decode_records, encode_records, export_state each record decoded through its copybook beside the raw bytes — the CP037 EBCDIC image and the packed/zoned nibbles exactly as stored, sign nibbles included; whole datasets, DB2-class tables and spool dumped as hex plus decoded rows with checksums, so "expected result" can be a byte, not a screenshot.

The workflow, end to end

1 · Seed the synthetic samples
Programs, copybooks and JCL go in as members; sample data goes in as records through the copybook (encode_records) or as EBCDIC datasets. Everything lands in a private workspace — a snapshot you can reset_workspace back to between runs.
2 · Analyse before touching anything
inventory_scan, find_program, explain_program and the analysis_* tools return the call graph, copybook usage, impact set and metrics as data an agent can reason over — the "understand the estate" step that every modernization guide puts first, done against the artifacts rather than from memory.
3 · Execute the workflow
Compile, run the batch, drive the transactions, read the VSAM clusters. One tool call per step, every step leaving a joblog, a spool file or a state export behind.
4 · Compare against expected results
Your expected results — captured from your licensed system by your own staff — are compared with the returned bytes and reports on your side. If you would rather we run the comparison, that is the migration-verification engagement: golden runs, byte-level diff, divergences explained with evidence or listed as unexplained.

Honesty line: CobolStack environments are independent clean-room implementations — specification-conformant, never "identical". Layouts, encodings, arithmetic, JCL semantics, batch chains and return codes are valid here; real-system timing, locking minutiae and abend behaviour beyond documented semantics are not. Your licensed system remains the sole conformance authority, and we put that sentence in every proposal.

//DEMAND  DD DSN=WHAT.THE.MARKET.ASKS.FOR,DISP=SHR

What the market is asking for in 2026 — and where we sit

Read from the public record, with sources at the foot of the page. We do not quote a market-size figure for "mainframe MCP services" because we could not find one we trust; if a vendor gives you one, ask for its method.

TrendEvidenceWhere CobolStack fits
Agents inspecting live z/OS through MCP The Zowe project ships an MCP server that lets agents list datasets, view members, submit jobs and issue TSO and console commands, with tools classified by risk so side-effect operations require confirmation[1][2]. IBM's own coverage describes agents that read JCL or COBOL, explain it and propose modernization paths as the emerging capability[1]. Complementary, not competing. Zowe-side tooling inspects the real estate read-only; the CobolStack servers are where an agent can execute — compile, submit, drive a transaction, mutate a VSAM cluster — without a change ticket against licensed iron.
IBM i joining the agent ecosystem IBM published an open-source IBM i MCP Server that exposes Db2 for i SQL queries as agent tools, fronted by the ContextForge gateway for access governance and audit[3]. ARCAD offers an MCP server with 70+ tools giving agents impact analysis and cross-reference context from its repository[4]. IBM markets Bob for IBM i modernization across RPG, CL, SQL and DDS[5]. SteelFrame X is the executable half: RPG IV/RPG III, CL, DDS and COBOL compile and run; 5250 screens drive; job queues, spool and journal behave — so the agent that ARCAD or Bob informed can actually rehearse the change.
Explain-and-validate for COBOL and JCL IBM's watsonx Code Assistant for Z positions generative AI to understand, explain, refactor, optimize, transform and validate COBOL, and adds natural-language JCL explanation[6]. Explanation needs something to be checked against. explain_program and the analysis tools give an agent the dependency facts; run_batch and export_state give it the outcome to test its explanation with.
Governed agents in mainframe operations BMC added MCP capabilities to its AMI Assistant and a Control-M MCP server so agents can interact with production workflows "within enterprise-grade governance, visibility, and policy controls"[7]. Same governance instinct, applied to rehearsal: scoped tokens, authority ceilings, single-tenant instances, audit on every tool call — and a RAM-only private lane that retains nothing.
The correction: verification is the bottleneck Gartner predicts more than 70% of mainframe exit projects started in 2026 will fail to deliver their intended benefits because of overestimated generative-AI capability, and suggests GenAI is often better used to modernize in place[8][9]. Practitioners keep reporting the same failure sites: packed-decimal rounding, EBCDIC-to-ASCII mismatches, REDEFINES overlay semantics[10]. This is the page's whole premise. Whoever writes the code, the proof is a byte-level comparison of what ran. The servers exist to make that comparison cheap enough to run on every iteration.
//SVCS    DD DSN=SERVICES.AROUND.THE.SERVERS

Services around the servers

Fixed scope first; a written capability statement attached to every proposal. We never log on to your production systems.

EngagementDeliverableNotes
Evaluation sprintyour synthetic samples seeded, the five validation items above exercised end to end by your agent (or ours), results returned as joblogs, spool, state exports and screen transcripts — plus a short written finding on what matched, what did not, and whythe fastest way to find out whether the platform fits your workflow; scoped in days
Agent-loop enablementyour coding agent or LLM tooling wired to the servers: token issuance with the right scopes, workspace conventions, prompt and tool-description review, the compile → run → decode → compare loop scripted as a reusable harnessworks with any MCP-capable client; we do not resell agent products
Modernization-validation harnessthe fixture-capture protocol your staff run on the licensed system, the expected-result store, and the comparison that turns the servers' state exports into a pass/fail with evidencethe same method as migration verification, packaged for agent-driven iteration
Estate analysis packinventory, call graph, copybook usage, impact sets and metrics exported from analysis_* and find_program for your artifacts — as data your agents and your architects both readpairs with the seam assessment where the estate crosses platforms
Governance setupscoped credentials, read-only versus execute tiers, single-tenant instance, audit routing, the ephemeral lane for zero-retention workenabled per organisation under your security review
Shop-standard estate for agentsa golden environment built to your naming, PROCs, security model and calendars so the agent rehearses in your shop's shape, not a generic onesee shop-standard image

What the servers are

  • An executable reference for the mainframe and IBM i sides — real compile, real job streams, real spool, real records
  • An oracle that returns evidence: bytes, listings, joblogs, transcripts, checksums
  • Isolated per workspace; persistent or RAM-only; every call authorised and audited

What they are not

  • Not a code converter and not a judge of equivalence — the comparison is a separate, visible step
  • Not a connection to your licensed system; nothing here touches production
  • Not self-serve: access is issued per organisation, with scopes, after a conversation

Sources: [1] Open Mainframe Project, "Hello, Mainframe: Using Zowe in an AI World with MCP Agents", 1 Jul 2026 · [2] zowe/zowe-mcp on GitHub (EPL-2.0) · [3] IBM Community, "Bridging IBM i and Modern AI — IBM i MCP Server and ContextForge", 14 Jul 2026 · [4] ARCAD Software, "ARCAD MCP Server — IBM i Application Context for AI Agents" · [5] IBM, "IBM i Modernization with IBM Bob" (product page) · [6] IBM, "What's new in watsonx Code Assistant for Z 2.1" · [7] BMC Software press release, 14 Jul 2026 · [8] Gartner press release, 18 Jun 2026 · [9] CIO Dive, "Overestimating AI threatens legacy mainframe migrations", 22 Jun 2026 · [10] DEV Community, "Mainframe Migration Tools: What Works and What Fails" (practitioner write-up). Vendor claims are reported as made; none implies endorsement of CobolStack. Sources checked September 2026.