Skip to content
Cloud InfrastructureQuality Score: 46/99 (Fair)No Auth RequiredSpec v2017-07-25auto GenerationTransport: stdio

AWS Glue DataBrewMCP Configuration & Schema Registry

The AWS Glue DataBrew Model Context Protocol (MCP) configuration provides a validated, machine-readable JSON schema and executable bridge that connects state-of-the-art AI coding assistants — including Claude Desktop, Cursor IDE, Windsurf, Cline, and VS Code Copilot — directly to the AWS Glue DataBrew REST API. By leveraging the standardized open Model Context Protocol, AI agents can dynamically discover capabilities, validate input parameters against strict JSON Schemas, and execute live API operations without context switching or manual copy-pasting.

Quick Specs & Integration Summary

1. Functionality:Exposes 10 API endpoints as callable AI tools for AWS Glue DataBrew.
2. Authentication:Zero authentication required — ready for immediate execution.
3. Protocol Layer:Standard Model Context Protocol JSON-RPC 2.0 via stdio transport.
4. Quick Launch:npx -y @modelcontextprotocol/server-openapi https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json

Technical Architecture & Protocol Semantics

Under the Model Context Protocol specification, the AWS Glue DataBrew configuration functions as an isolated protocol adapter. When an AI agent initializes a session, the client establishes a bidirectional JSON-RPC 2.0 communication channel over standard input/output (stdio) or Server-Sent Events (SSE). During the initial handshake, the server publishes its tool manifest extracted from the AWS Glue DataBrew OpenAPI specification (version 2017-07-25).

The AWS Glue DataBrew API, provided by Amazon Web Services, serves as the programmatic backbone for DataBrew, a fully managed, visual data preparation service designed to accelerate data processing for analytics and machine learning. This API exposes the core functionalities of the service, enabling developers to programmatically create, manage, and execute data preparation workflows. Its primary capabilities include orchestrating dataset profiling jobs to uncover data quality issues, managing reusable recipe versions that contain data cleansing and transformation steps, and triggering batch or interactive jobs to apply these recipes at scale. The API is fundamental for enterprise data engineering teams, data scientists, and analysts who need to automate data pipelines, enforce data quality standards, and prepare vast, complex datasets stored in Amazon S3 or connected data stores without writing extensive ETL code. Typical use cases range from automating the cleansing of incoming IoT sensor data for predictive maintenance to standardizing disparate customer data sources for a unified analytics platform. Exposing the AWS Glue DataBrew API through tools like the Model Context Protocol (MCP) transforms it into a dynamic, interactive resource for an AI coding assistant. This integration allows the AI to act as a collaborative data preparation engineer, directly interfacing with the data lifecycle within a developer's cloud environment. Instead of merely generating boilerplate code, the assistant can perform actionable operations such as querying the current state of datasets or recipes, analyzing the output of a profile job to identify specific data quality anomalies, and then initiating targeted recipe steps to address those issues. The value lies in bridging the gap between high-level intent and executable cloud infrastructure actions; a developer can describe a data problem, and the AI, via MCP, can investigate the live environment, propose a DataBrew-based solution, and even implement it by making precise API calls, thereby drastically reducing context-switching and accelerating iteration cycles. Within an MCP-enabled environment, a developer can instruct the AI agent to perform a range of dynamic, context-aware tasks using the DataBrew API. For instance, an agent can be directed to "generate a new dataset from the S3 path 's3://company-data/sales-2024/raw/' and run a profile job, then summarize the key statistics and data quality findings." Following analysis, the AI could then be instructed to "create a new recipe version to fix the identified missing values in the 'customer_id' column and apply a standardization transformation to the 'product_code' field." Furthermore, the agent can automate repetitive maintenance by being told to "list all active recipes, check their last run status, and trigger a re-run of any recipe that has failed in the past 24 hours." This transforms the AI from a code generator into an operational assistant capable of monitoring, analyzing, and remediating data pipelines through direct, secure API interaction. Critical to the implementation of an MCP server for the DataBrew API is the strict adherence to security and authentication best practices, despite the mention of "None" in the endpoint list, which refers to the API's own scheme, not the server's authentication. In practice, the MCP server itself must be secured. The most fundamental requirement is configuring the server with robust AWS IAM credentials that possess only the necessary DataBrew permissions, following the principle of least privilege. A dedicated IAM role or user should be created with policies granting minimal access, such as `databrew:ListDatasets` and `databrew:GetDataset` for read-only monitoring, or specific job and recipe action permissions for operational tasks, rather than broad administrative rights. Developers must ensure these credentials are never exposed and are managed via secure environment variables or a secrets manager. Additionally, enabling AWS CloudTrail logging for DataBrew API calls provides a vital audit trail for all actions performed by the AI agent through the MCP server. This architecture guarantees strict process boundary isolation: all sensitive authorization headers and secret tokens remain sandboxed inside the client runtime, never leaking into language model context windows or external logging endpoints.

Authentication TypePublic (No Auth)Injected via local client environment
Tools & Routes Mapped10 OperationsConforms to JSON-RPC 2.0 specs
Specification OriginOpenAPI v2017-07-25auto schema validation
Documentation & Schema Quality Index
46
★ Grade C - Baseline Coverage
Automated Audit Checklist
Automated schema extraction & validation (+12 pts)
Extensive tool mapping (10 endpoints defined) (+20 pts)
Zero-configuration public API instant execution (+20 pts)
Full JSON-RPC 2.0 Model Context Protocol specification conformity (+15 pts)
Upstream technical documentation verification (+12 pts)

Hosted Remote Configuration URL

MCP Configuration File

Provide this hosted URL in any client that supports remote MCP schema auto-loading.

https://mcpbridge.org/config/amazonaws-com-databrew.json

2. AI Assistant Use Cases & Practical Workflows

Tailored for Cloud Infrastructure

Real-world execution scenarios demonstrating how LLM agents (Claude 3.7, GPT-4o, Cursor Agent) invoke AWS Glue DataBrew tools to automate developer workflows.

1. CI/CD Build Failure & Telemetry Diagnostics

CI/CD Remediation

Instantly diagnose failing CI/CD builds or deployment pipelines by streaming build logs, isolating failure root causes, and drafting targeted code fixes.

Example Natural Language Prompt:

"Fetch recent pipeline run logs from AWS Glue DataBrew. Isolate the failed step, summarize the exact compiler or test failure error, and propose a pull request fix in Cursor."

Mapped: /recipes/{name}/batchDeleteRecipeVersion

2. Cloud Resource Auditing & Cost Optimization

Cloud FinOps

Scan active compute clusters, storage buckets, and networking configurations to identify unattached volumes or idle oversized instances.

Example Natural Language Prompt:

"Query active cloud infrastructure resources in AWS Glue DataBrew. Identify unattached storage volumes, idle compute instances, and summarize estimated monthly cost savings."

Mapped: /datasets

3. Zero-Downtime Rollout & Canary Health Verification

Deployment Ops

Orchestrate progressive deployments, monitor error rate thresholds on newly deployed pods, and execute automated rollbacks if error budgets breach.

Example Natural Language Prompt:

"Check the active deployment rollout status in AWS Glue DataBrew. Monitor canary error rate percentages for 5 minutes and report whether the deployment is safe to promote to 100% traffic."

Autonomous Agent Loop

4. Infrastructure as Code (IaC) Drift Detection

IaC Governance

Compare live deployed resource state against Terraform or CloudFormation definitions to spot unauthorized manual changes.

Example Natural Language Prompt:

"Scan live configurations via AWS Glue DataBrew and compare against our repository IaC definitions. Highlight any configuration drift in security groups or network routes."

Autonomous Agent Loop

End-to-End Multi-Step Agent Execution Lifecycle

When an engineer submits a task to Claude Desktop or Cursor, the LLM executes an autonomous 4-phase Model Context Protocol loop:

Phase 1

Schema Introspection

Handshake lists all 10 tools and builds argument validators.

Phase 2

Argument Synthesis

Model extracts parameters from prompt and validates types against OpenAPI rules.

Phase 3

Stdio Execution

Bridge invokes live API with injected local credentials and captures raw HTTP response.

Phase 4

Output Remediation

LLM parses JSON results, handles status codes, and presents synthesized answers.

3. Multi-Client Installation Matrix & Setup Guides

Select your AI assistant below to view exact configuration file paths, JSON installation snippets, and launch commands.

Claude Desktop

claude_desktop_config.json
macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
Windows: %APPDATA%\Claude\claude_desktop_config.json
Linux: ~/.config/Claude/claude_desktop_config.json
{
  "mcpServers": {
    "amazonaws-com-databrew": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/server-openapi",
        "https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"
      ],
      "env": {
        "AWS_GLUE_DATABREW_API_KEY": "your_aws_glue_databrew_api_key"
      }
    }
  }
}
Deep link

Cursor IDE

.cursor/mcp.json

Open Cursor Settings → Features → MCP Servers, or create .cursor/mcp.json in your project root.

{
  "mcpServers": {
    "amazonaws-com-databrew": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/server-openapi",
        "https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"
      ],
      "env": {
        "AWS_GLUE_DATABREW_API_KEY": "your_aws_glue_databrew_api_key"
      }
    }
  }
}

Saves as .cursor/mcp.json in the download. Move it to your project root.

Deep link install →

VS Code / Cline Extension

cline_mcp_settings.json

Paste into your Cline extension MCP configuration or Roo Code host settings.

{
  "mcpServers": {
    "amazonaws-com-databrew": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/server-openapi",
        "https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"
      ],
      "env": {
        "AWS_GLUE_DATABREW_API_KEY": "your_aws_glue_databrew_api_key"
      }
    }
  }
}

Zed Editor & Docker CLI

Zed / Docker

Docker container execution command:

docker run -i --rm -e AWS_GLUE_DATABREW_API_KEY="YOUR_SECRET_VALUE" node:20-alpine npx -y @modelcontextprotocol/server-openapi https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json

Zed settings context servers JSON:

{
  "context_servers": {
    "amazonaws-com-databrew": {
      "command": {
        "path": "npx",
        "args": [
          "-y",
          "@modelcontextprotocol/server-openapi",
          "https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"
        ],
        "env": {
          "AWS_GLUE_DATABREW_API_KEY": "your_aws_glue_databrew_api_key"
        }
      }
    }
  }
}

Programmatic SDK Integration (TypeScript / Python)

Initialize the AWS Glue DataBrew MCP client directly in your backend codebase.

import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";

// Initialize AWS Glue DataBrew MCP client transport over stdio
const transport = new StdioClientTransport({
  command: "npx",
  args: ["-y","@modelcontextprotocol/server-openapi","https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"],
  env: { AWS_GLUE_DATABREW_API_KEY: process.env.AWS_GLUE_DATABREW_API_KEY || "YOUR_SECRET_KEY" }
});

const client = new Client(
  { name: "amazonaws-com-databrew-client", version: "1.0.0" },
  { capabilities: { tools: {}, resources: {}, prompts: {} } }
);

async function connectAndRun() {
  await client.connect(transport);
  const tools = await client.listTools();
  console.log("Connected to AWS Glue DataBrew MCP Server.");
  console.log("Discovered 10 mapped tools:", tools);
}

connectAndRun().catch(console.error);

Raw Stdio Schema Definition

schema.json

For standalone CLI wrappers, background daemon daemons, or custom script integrations:

{
  "mcpServers": {
    "amazonaws-com-databrew": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/server-openapi",
        "https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json"
      ],
      "env": {
        "AWS_GLUE_DATABREW_API_KEY": "your_aws_glue_databrew_api_key"
      }
    }
  }
}

4. Security, Authentication & Credential Management

Safely configure authentication tokens, isolate execution environments, and implement enterprise security best practices.

Required Environment Keys Reference

Variable NameRequiredTypeDefaultPurpose & Guidance
AWS_GLUE_DATABREW_API_KEYREQUIREDSecret Key / TokenNone (Set in env)your_aws_glue_databrew_api_key

Zero-Downtime Token Rotation Protocol

  1. Generate Secondary Key: Create a new secret API token with identical scopes in your AWS Glue DataBrew developer portal.
  2. Update Client Configuration: Insert the new token inside the env block of your MCP client JSON config.
  3. Validate Connection: Issue a test query in Claude or Cursor to ensure handshake and tool calls succeed.
  4. Revoke Stale Token: Decommission the legacy key on the vendor portal to prevent unauthorized access.

Least-Privilege & Sandboxing Rules

  • Read-Only Token Scoping: Whenever your workflow only requires querying data, provision read-only credentials to prevent accidental mutations.
  • Local Process Isolation: Stdio transports run in isolated local subprocesses; secret credentials are never sent across the internet to MCP Bridge servers.
  • Prompt Injection Defense: AI model responses are sandboxed; verify generated destructive arguments before confirming execution in agent mode.

Enterprise Security Checklist (Mandatory Practices)

  • Never commit claude_desktop_config.json or .cursor/mcp.json containing raw secrets into public GitHub repositories.
  • Add .cursor/mcp.json and .env.local to your project's .gitignore file.
  • Always enforce TLS/HTTPS encryption on outbound network requests initiated by the server process.

5. Tool Parameter Schemas & Natural Language Execution

Mapped OpenAPI operations converted into discrete Model Context Protocol tools with strict JSON-RPC payload validators.

10 Total Tools Mapped
POST/recipes/{name}/batchDeleteRecipeVersion
tools/call: amazonaws-com-databrew_post_recipes__name__batchDeleteRecipeVersion

BatchDeleteRecipeVersion

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_post_recipes__name__batchDeleteRecipeVersion",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute BatchDeleteRecipeVersion and output the formatted result."

GET/datasets
tools/call: amazonaws-com-databrew_get_datasets

ListDatasets

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_get_datasets",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute ListDatasets and output the formatted result."

POST/datasets
tools/call: amazonaws-com-databrew_post_datasets

CreateDataset

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 3,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_post_datasets",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute CreateDataset and output the formatted result."

POST/profileJobs
tools/call: amazonaws-com-databrew_post_profileJobs

CreateProfileJob

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 4,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_post_profileJobs",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute CreateProfileJob and output the formatted result."

GET/projects
tools/call: amazonaws-com-databrew_get_projects

ListProjects

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 5,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_get_projects",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute ListProjects and output the formatted result."

POST/projects
tools/call: amazonaws-com-databrew_post_projects

CreateProject

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 6,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_post_projects",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute CreateProject and output the formatted result."

GET/recipes
tools/call: amazonaws-com-databrew_get_recipes

ListRecipes

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 7,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_get_recipes",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute ListRecipes and output the formatted result."

POST/recipes
tools/call: amazonaws-com-databrew_post_recipes

CreateRecipe

Zero required query/path parameters for this endpoint.
JSON-RPC 2.0 Request Payload
{
  "jsonrpc": "2.0",
  "id": 8,
  "method": "tools/call",
  "params": {
    "name": "amazonaws-com-databrew_post_recipes",
    "arguments": {}
  }
}
Natural Language Prompt

"Use AWS Glue DataBrew to execute CreateRecipe and output the formatted result."

6. Interactive Troubleshooting & FAQ Accordion

Diagnose and resolve common JSON-RPC protocol error codes, connection disconnects, and schema refresh issues.

A 401 Unauthorized response indicates that the upstream AWS Glue DataBrew API rejected the authentication credential supplied in your MCP client's environment configuration. To resolve this: (1) Verify that your secret token is defined inside the "env" block of claude_desktop_config.json or .cursor/mcp.json rather than hardcoded in the command string. (2) Check whether AWS Glue DataBrew requires a prefix such as "Bearer <token>" in the authorization header. (3) Confirm that your API key has not expired and has been granted sufficient least-privilege scopes on the AWS Glue DataBrew developer dashboard.

If your MCP client fails to initialize tools for AWS Glue DataBrew: (1) Test the bridge launcher command ("npx -y @modelcontextprotocol/server-openapi https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json") directly inside your terminal or shell to inspect stdout/stderr diagnostic traces. (2) Verify network connectivity to the schema source (https://api.apis.guru/v2/specs/amazonaws.com/databrew/2017-07-25/openapi.json). (3) Ensure Node.js (v18+) is installed and accessible in your system PATH. (4) For authenticated APIs, confirm credentials are configured in your client's "env" mapping rather than command arguments.

Similar Cloud Infrastructure Configurations

Explore related API bridges with ready-to-use Model Context Protocol schemas.

Supabase API

Cloud Infrastructure

Manage Supabase projects, databases, authentication, and storage through your AI agent.

https://mcpbridge.org/config/supabase.json

Cloudflare API

Cloud Infrastructure

Manage Cloudflare DNS, CDN, Workers, and security settings through your AI agent.

https://mcpbridge.org/config/cloudflare.json

Vercel API

Cloud Infrastructure

Deploy projects, manage domains, and monitor deployments through your AI agent.

https://mcpbridge.org/config/vercel.json

DigitalOcean API

Cloud Infrastructure

The DigitalOcean API is a comprehensive, RESTful interface provided by DigitalOcean, a leading cloud infrastructure provider focused on simplifying cloud computing for developers, startups, and enterprises. It serves as the programmatic backbone for managing the entire DigitalOcean ecosystem, enabling users to provision, configure, and control cloud resources such as Droplets (virtual private servers), Kubernetes clusters, managed databases, networks, storage volumes, and application platforms. Core capabilities include full lifecycle management of these resources, from creation and scaling to monitoring and deletion, mirroring the functionality available in the DigitalOcean control panel. Its primary use cases range from automating infrastructure setup for CI/CD pipelines and enabling infrastructure-as-code practices to supporting dynamic application scaling and resource optimization for SaaS products, e-commerce sites, and development environments. The API is designed for both developers seeking to automate their cloud operations and businesses that require programmable, scalable cloud infrastructure without the complexity of larger hyperscale providers. When exposed as tools via the Model Context Protocol (MCP) to an AI coding assistant, the DigitalOcean API transforms from a traditional developer tool into a dynamic, context-aware resource for intelligent infrastructure automation. The MCP server acts as a bridge, allowing the AI model to understand and execute API calls based on natural language instructions and the current project context. This integration provides immense value by enabling the AI to perform real-time cloud management tasks directly within the development workflow. For instance, the AI can instantly query account details to verify resources, list and manage SSH keys for secure access, or retrieve and monitor the status of infrastructure actions. This contextual access means the AI can make informed suggestions or take automated actions—like recommending a cost-optimized Droplet size based on current usage patterns or verifying that a new SSH key has been correctly added before proceeding with a deployment script—thereby reducing context-switching and accelerating development cycles. Practical workflow examples demonstrate the power of this MCP integration. A developer could instruct the AI agent with commands like, "Query our account for all active SSH keys and ensure the one named 'ci-bot' is present; if not, create it using this public key," automating a common security and setup step. Another example involves asking the AI to "Check the status of our last ten infrastructure actions to see if any are stuck in a 'pending' state," which would leverage the actions endpoints to provide an immediate operational health check. More complex automations are possible, such as "Based on the current Droplet inventory from the API, generate a Terraform configuration file that replicates this setup," or "Scan our Kubernetes 1-Click apps and suggest one for deploying a new microservice based on the project requirements." These interactions turn the AI into a proactive DevOps partner capable of auditing, reporting, and modifying cloud infrastructure through simple, conversational directives. Critical to the secure operation of this MCP server is rigorous attention to authentication and access control, despite any initial configuration notes indicating "None" for simplicity. In any real-world deployment, authentication via a DigitalOcean Personal Access Token is non-negotiable. This token should be treated as a high-privilege secret. Developers must adhere to the principle of least privilege by creating tokens with the minimum scopes required for the specific tasks—such as read-only access for monitoring or write access only for specific resource types. Best practices include storing tokens in secure environment variables or a secrets manager, never hardcoding them, and ensuring the MCP server configuration does not expose them in logs or client-side code. Furthermore, regular token rotation and monitoring of API activity through DigitalOcean's audit logs are essential to maintain a secure posture when integrating cloud management capabilities directly into AI-assisted development environments.

https://mcpbridge.org/config/digitalocean-com.json