Introduction
Most CTF (Capture The Flag) platforms heavily prioritize offensive mechanics-finding vulnerabilities and extracting flags. But does this model actually help developers understand the root cause of security bugs? More importantly, how effectively does it bridge the gap between offensive exploitation and secure coding practices?
Driven by these questions, I set out to build RedPatch : an open-source, hybrid application security (AppSec) playground engineered for both developers and security researchers. RedPatch adopts a dual approach combining Red Team (offensive) and Blue Team (defensive) workflows. Rather than focusing solely on exploitation, it forces you to analyze the underlying source code and write secure patches.
By integrating Large Language Models (LLMs), RedPatch automatically evaluates whether a submitted patch successfully mitigates the security flaw, serving real-time, actionable feedback directly to the user.
System Architecture
Project Structure
To maintain a clean, modular architecture and encourage community contributions, the core orchestration engine is strictly decoupled from individual lab environments via isolated Docker containers across separate repositories.
redpatch/
├── app/
│ ├── main.py
│ │
│ ├── core/
│ │ ├── config.py
│ │ └── config.json
│ │
│ ├── labs/
│ │ └── manifest.json
│ │
│ ├── services/
│ │ ├── ai/
│ │ ├── container_services/
│ │ └── module_manager/
│ │
│ ├── static/
│ └── templates/
│
├── docker-compose.yaml
├── requirements.txt
├── CONTRIBUTING.md
├── SECURITY.md
└── LICENSE
Grounded in this decoupled design, external developers can seamlessly register custom lab environments using the project’s manifest.json specification:
{
"labs": {
"SQLi": {
"description": "SQL Injection",
"submodules": [
{
"id": "sqli-0",
"title": "SQL Injection - Authentication Bypass",
"category": "web",
"image_tag": "redpatch-lab/sqli-0:v1.0.0",
"port": 5000,
"dev_path": "./labs/sqli-0",
"download_url": "....tar.gz"
}
]
}
}
}
Tech Stack Breakdown
| Component | Technology | Technical Role |
|---|---|---|
| Backend Engine | FastAPI (Python) | High-performance asynchronous API, routing, and core orchestration |
| Lab Isolation | Docker Desktop / Engine | Isolated containerized execution environments for vulnerable targets |
| AI Auditor | Google Gemini API | Automated security auditing and patch validation engine |
| Frontend UI | Jinja2 / HTML5 / Tailwind | UI layout, integrated Monaco code editor, and interactive terminal interface |
To prevent tight coupling with a single AI ecosystem, the LLM provider engine is abstractly interfaced through a RedTeamAgent class. The active provider instance is dynamically returned at runtime based on the config.json parameters:
class RedTeamAgent:
def __init__ (self):
self.provider: BaseLLMProvider = self._get_provider()
def _get_provider(self) -> BaseLLMProvider:
provider_name = settings.LLM_PROVIDER.lower()
if provider_name == "gemini":
return GeminiProvider()
else:
raise ValueError(f"Unsupported or undefined LLM provider: {provider_name}")
async def run_attack(self, code: str, vulnerability_type: str, routes: list, lab_link: str) -> VulnerabilityAnalysis:
return await self.provider.analyze_code(code, vulnerability_type, routes, lab_link)
Technical Details & Core Mechanics
1. Containerized Workspace Isolation
Every lab environment spawns as an independent Docker container separate from the primary engine. Real-time code modifications applied within the embedded editor are dynamically mounted into a temporary workspace volume inside the target container.
2. Dual-Mode Workflow
- Pentester Mode: Uncover vulnerabilities, construct payloads, execute exploits, and retrieve flags.
- Coder Mode: Inspect raw source code within the embedded Monaco Editor, diagnose structural flaws, and refactor code to enforce secure coding practices.
3. Automated Patch Verification Engine
At runtime, RedPatch inspects the target application’s route structures and feeds the entire modified source code — along with the active target URL and vulnerability classification — to the Gemini API. By leveraging Gemini’s native response_schema feature, the engine guarantees strictly typed structural output adhering to a JSON specification:
{
"type": "object",
"properties": {
"vulnerability_found": {"type": "boolean"},
"target_line": {"type": "integer"},
"explanation": {"type": "string"},
"exploit_request": {
"type": "object",
"properties": {
"path": {"type": "string", "description": "The HTTP endpoint path, e.g., /login-vulnerable"},
"method": {"type": "string", "description": "HTTP Method in uppercase: POST, GET, PUT, DELETE"},
"headers": {"type": "string", "description": "JSON string of headers or empty string"},
"params": {"type": "string", "description": "URL query string or empty string"},
"data": {"type": "string",
"description": "Form payload string, e.g. username=admin&password=123, or empty string"},
"json_body": {"type": "string", "description": "JSON body string or empty string"}
},
"required": ["path", "method"]
}
},
"required": ["vulnerability_found", "target_line", "explanation", "exploit_request"]
}
This structural output allows users to fire the AI-generated attack vector against their modified target with a single click in the UI. If the patch successfully mitigates the attack vector, the AI acknowledges the remediation and flags the module as solved.
Engineering Challenges & Lessons Learned
Architecting RedPatch involved navigating several non-trivial system design hurdles:
1. Decoupled Docker Container Management
RedPatch was my first major project built with Docker, and I prioritized two primary constraints:
- Complete process and network isolation for target lab instances.
- Effortless environment setup by allowing the core platform itself to deploy within a container.
A primary engineering challenge was handling workspace syncing across container boundaries. Transmitting code changes over HTTP endpoints introduced unnecessary overhead and failure points. Instead, I implemented a temporary host workspace mounted directly into the target lab container. This guaranteed hot-reloading when code edits occurred in the web editor.
Additionally, to accommodate execution environments both inside and outside Docker containers, I added an ENV IS_DOCKER=true variable within the Dockerfile. A runtime helper function checks this flag to construct valid target workspace paths across environment configurations.
2. Enforcing Deterministic AI Outputs
In my initial implementation, I passed Pydantic models directly into the Gemini API’s native response_schema parameter. However, the model struggled to interpret the constraints of complex Pydantic schemas accurately—it frequently misconstrued critical fields as optional and omitted them, causing constant verification failures during downstream parsing.
To resolve this, I refactored the pipeline by feeding a simplified, explicit JSON Schema directly into response_schema at the API level. This ensured strict structural adherence directly during the model's generation stage. I then passed the resulting raw JSON payload back to Pydantic purely for the final validation pass.
Getting the model to generate accurate exploit HTTP requests was another hurdle. It required heavy prompt tuning, but dynamically injecting extracted application routes and signature parameters into the context window finally resolved payload inconsistencies.
Future Roadmap & Conclusion
Looking ahead, planned updates for RedPatch include:
- Expanding the vulnerable lab inventory to cover broader OWASP Top 10 classifications.
- Patching identified bugs and refining system execution pipelines.
- Introducing multi-user capabilities and competitive dynamic scoring leaderboards.
Contributing & Supporting the Project
If you are interested in contributing to RedPatch:
- Check out CONTRIBUTING.md, build custom lab modules, and submit a Pull Request to the labs repository.
- Propose feature enhancements or report system bugs via GitHub Issues.
- Support project development by dropping a star on GitHub! ⭐
GitHub Repositories:
- RedPatch Core Engine: https://github.com/msalihberk/redpatch
- RedPatch Labs Inventory: https://github.com/msalihberk/redpatch-labs
📌 References & Community
If you want to check out my other security research, tools, or open-source projects, feel free to explore the links below:
- GitHub: github.com/msalihberk
- X: x.com/msalihberkk
- Previous Research: Is the Android Lock Screen an Illusion? A Critical Logical Bypass Discovered in the Gemini App
- Follow for More: Feel free to follow my Medium profile to get notified about my future security research, development projects, and technical write-ups.






![FastAPI's detail has two shapes, and the second one shows your users [object Object]](https://media2.dev.to/dynamic/image/width=1000,height=420,fit=cover,gravity=auto,format=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl5tchh7xlof5a6lgp5kg.png)








