An open-source autonomous browser agent that can navigate websites, interact with web applications, and execute multi-step tasks using AI. This is the core browser automation engine that powers TaskCato.com - the AI co-founder platform for founders and startups.
This is a headless browser agent controlled by AI (LLM) that can:
- Navigate websites autonomously
- Fill out forms and interact with web UIs
- Extract data and take screenshots
- Execute multi-step workflows
- Remember context across sessions
- Run multiple browser instances simultaneously
Think of it as "Selenium meets ChatGPT" - but the AI decides what to do next.
| Feature | Open Source (This Repo) | TaskCato.com (SaaS) |
|---|---|---|
| Browser Automation | ✅ Full access | ✅ Enhanced |
| AI-Powered Navigation | ✅ Yes | ✅ Yes |
| Multi-Instance Support | ✅ Up to 9 browsers | ✅ Unlimited |
| Integrations | ❌ None | ✅ 15+ (HubSpot, Gmail, etc.) |
| Voice Calls | ❌ No | ✅ AI phone calls |
| Persistent Memory | ✅ Basic | ✅ Advanced company context |
| Team Collaboration | ❌ No | ✅ Yes |
| Support | Community | Priority |
| Hosting | Self-hosted | Cloud-hosted |
| Price | Free | $29/month |
Use this open-source version if:
- You want to build your own automation tools
- You need browser automation for personal projects
- You're a developer who wants to customize the AI behavior
- You prefer self-hosting
Use TaskCato.com if:
- You're a founder who needs end-to-end business automation
- You want integrations with CRM, email, calendar, etc.
- You need AI that can make phone calls and manage your entire workflow
- You want a managed service with support
We've made it dead simple. No Python, no terminal commands, no configuration.
- Download
Browser Agent.appfrom Releases - Double-click to launch
- That's it! The app handles everything automatically:
- Creates Python virtual environment
- Installs all dependencies
- Starts the server
- Opens the dashboard
- Download
Browser Agent.exefrom Releases - Double-click to launch
- Done! Same automatic setup as Mac
- The app will show a loading screen while it sets up (first time only)
- Dashboard opens automatically at
http://localhost:8000 - Click Settings to configure your LLM:
- LM Studio (recommended for privacy): Download from https://lmstudio.ai
- Gemini (cloud): Add your API key from https://aistudio.google.com
- Start giving commands!
"Research the top 5 AI startups and take screenshots"
"Go to Hacker News and summarize the top 3 stories"
"Fill out the contact form on example.com"
Want to modify the code or run from source?
- Python 3.13+
- Node.js 18+ (for Electron)
- LM Studio or Gemini API key
# Clone the repository
git clone https://github.com/CloudCorpRecords/browserboi.git
cd browserboi
# Option 1: Run the Electron app (recommended)
cd launcher
npm install
npm start
# Option 2: Run server only (terminal-based)
python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
playwright install chromium
python -m uvicorn browser_agent.server.main:app --host 0.0.0.0 --port 8000cd launcher
npm install
npm run build
# Outputs:
# Mac: dist/mac-arm64/Browser Agent.app
# Windows: dist/Browser Agent 1.0.0.exe┌─────────────────────────────────────────┐
│ Electron App (Optional UI) │
│ - Dashboard │
│ - 9-Grid Browser View │
│ - Real-time Logs │
└──────────────┬──────────────────────────┘
│
┌──────────────▼──────────────────────────┐
│ FastAPI Server (Port 8000) │
│ - WebSocket for real-time updates │
│ - Multi-instance management (1-9) │
│ - RESTful API │
└──────────────┬──────────────────────────┘
│
┌──────────────▼──────────────────────────┐
│ AI Agent Core │
│ - LLM Integration (Gemini/LM Studio) │
│ - Task Planning & Execution │
│ - Memory Management │
└──────────────┬──────────────────────────┘
│
┌──────────────▼──────────────────────────┐
│ Browser Manager (Playwright) │
│ - Headless Chromium │
│ - Screenshot Capture │
│ - Page Interaction │
└─────────────────────────────────────────┘
- Autonomous Navigation: AI decides where to click, what to type
- Multi-Step Workflows: Chain complex actions together
- Context Awareness: Remembers previous actions and decisions
- Screenshot Analysis: AI "sees" the page and makes decisions
- Text Extraction: Reads page content for better understanding
- Session Persistence: Maintains cookies and login state
- 9-Instance Grid: Run up to 9 browser agents simultaneously
- Headless Mode: Runs in background without visible browser
- CDP Support: Chrome DevTools Protocol for advanced control
- Memory System: Stores user preferences and company context
- Streaming Logs: Real-time visibility into agent actions
browser_agent/
├── core/
│ ├── agent.py # Main AI agent logic
│ ├── browser.py # Playwright browser manager
│ ├── llm.py # LLM integration
│ └── memory.py # Context & memory management
├── server/
│ ├── main.py # FastAPI server
│ └── static/ # Web UI
└── tools/ # Agent tools (click, type, etc.)
launcher/ # Electron desktop app
├── main.js # Electron main process
├── index.html # Loading screen
└── package.json # Dependencies
Start a task:
curl -X POST http://localhost:8000/api/start/1 \
-H "Content-Type: application/json" \
-d '{"task":"Search for AI news on Hacker News"}'Stop an agent:
curl -X POST http://localhost:8000/api/stop/1Health check:
curl http://localhost:8000/api/diagnosticConnect to ws://localhost:8000/ws/1 for real-time updates:
status- Agent status changeslog- Action logsscreenshot- New screenshotserror- Error messages
We welcome contributions! This is an open-source project maintained by the TaskCato team.
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
- 🐛 Bug fixes and stability improvements
- 📝 Documentation and examples
- 🔌 New tool integrations
- 🧪 Test coverage
- 🎨 UI/UX improvements
"Go to ProductHunt and extract the top 10 products today""Fill out the contact form on example.com with my info""Research the pricing pages of Stripe, Square, and PayPal""Test the login flow on staging.myapp.com""Find email addresses of CTOs at Series A fintech startups"# LLM Provider
LLM_PROVIDER=lm_studio # or 'gemini'
GEMINI_API_KEY=your_key_here
# Browser Settings
HEADLESS=true # Run browser in headless mode
VIEWPORT_WIDTH=1280
VIEWPORT_HEIGHT=800
# Server
PORT=8000
HOST=0.0.0.0Settings are stored in browser_agent/data/settings.json:
{
"provider": "lm_studio",
"lm_studio_url": "http://localhost:1234/v1",
"lm_studio_model": "qwen/qwen3-vl-8b",
"gemini_api_key": "",
"gemini_model": "gemini-2.0-flash-exp"
}- No Data Collection: All data stays on your machine
- Local LLM Support: Use LM Studio for complete privacy
- Session Isolation: Each browser instance is isolated
- No Telemetry: We don't track usage
- UI Message Display: Agent responses may not appear in chat (check terminal logs)
- Browser View Content: Grid cells show blank instead of live browser content
- Windows Support: Electron app tested primarily on macOS
See Issues for full list.
MIT License - see LICENSE file for details.
This browser agent is the core technology behind TaskCato.com - an AI co-founder platform that helps founders automate their entire business workflow.
TaskCato adds:
- 15+ SaaS integrations (HubSpot, Gmail, Slack, etc.)
- AI voice calling capabilities
- Advanced persistent memory
- Team collaboration features
- Managed cloud hosting
- Priority support
Try TaskCato free: https://taskcato.com
- Discord: Join our community
- Twitter: @TaskCato
- Issues: GitHub Issues
- Discussions: GitHub Discussions
Built with:
- Playwright - Browser automation
- FastAPI - Web framework
- Electron - Desktop app
- Google Gemini - AI capabilities
- LM Studio - Local LLM support
Made with ❤️ by the TaskCato team
Building the future of autonomous work, one browser action at a time.