Skip to content

About

Local browser automation agent powered by LM Studio — fully local LLMs

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Browser AI Agent

An open-source autonomous browser agent that can navigate websites, interact with web applications, and execute multi-step tasks using AI. This is the core browser automation engine that powers TaskCato.com - the AI co-founder platform for founders and startups.

🎯 What Is This?

This is a headless browser agent controlled by AI (LLM) that can:

  • Navigate websites autonomously
  • Fill out forms and interact with web UIs
  • Extract data and take screenshots
  • Execute multi-step workflows
  • Remember context across sessions
  • Run multiple browser instances simultaneously

Think of it as "Selenium meets ChatGPT" - but the AI decides what to do next.

🆚 Open Source vs TaskCato Premium

Feature Open Source (This Repo) TaskCato.com (SaaS)
Browser Automation ✅ Full access ✅ Enhanced
AI-Powered Navigation ✅ Yes ✅ Yes
Multi-Instance Support ✅ Up to 9 browsers ✅ Unlimited
Integrations ❌ None ✅ 15+ (HubSpot, Gmail, etc.)
Voice Calls ❌ No ✅ AI phone calls
Persistent Memory ✅ Basic ✅ Advanced company context
Team Collaboration ❌ No ✅ Yes
Support Community Priority
Hosting Self-hosted Cloud-hosted
Price Free $29/month

Use this open-source version if:

  • You want to build your own automation tools
  • You need browser automation for personal projects
  • You're a developer who wants to customize the AI behavior
  • You prefer self-hosting

Use TaskCato.com if:

  • You're a founder who needs end-to-end business automation
  • You want integrations with CRM, email, calendar, etc.
  • You need AI that can make phone calls and manage your entire workflow
  • You want a managed service with support

🚀 Quick Start

One-Click Installation

We've made it dead simple. No Python, no terminal commands, no configuration.

Mac Users

  1. Download Browser Agent.app from Releases
  2. Double-click to launch
  3. That's it! The app handles everything automatically:
    • Creates Python virtual environment
    • Installs all dependencies
    • Starts the server
    • Opens the dashboard

Windows Users

  1. Download Browser Agent.exe from Releases
  2. Double-click to launch
  3. Done! Same automatic setup as Mac

First Run

  1. The app will show a loading screen while it sets up (first time only)
  2. Dashboard opens automatically at http://localhost:8000
  3. Click Settings to configure your LLM:
  4. Start giving commands!

Example Commands

"Research the top 5 AI startups and take screenshots"
"Go to Hacker News and summarize the top 3 stories"
"Fill out the contact form on example.com"

🛠️ For Developers

Want to modify the code or run from source?

Prerequisites

  • Python 3.13+
  • Node.js 18+ (for Electron)
  • LM Studio or Gemini API key

Running from Source

# Clone the repository
git clone https://github.com/CloudCorpRecords/browserboi.git
cd browserboi

# Option 1: Run the Electron app (recommended)
cd launcher
npm install
npm start

# Option 2: Run server only (terminal-based)
python3 -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate
pip install -r requirements.txt
playwright install chromium
python -m uvicorn browser_agent.server.main:app --host 0.0.0.0 --port 8000

Building the App

cd launcher
npm install
npm run build

# Outputs:
# Mac: dist/mac-arm64/Browser Agent.app
# Windows: dist/Browser Agent 1.0.0.exe

🏗️ Architecture

┌─────────────────────────────────────────┐
│  Electron App (Optional UI)             │
│  - Dashboard                             │
│  - 9-Grid Browser View                   │
│  - Real-time Logs                        │
└──────────────┬──────────────────────────┘
               │
┌──────────────▼──────────────────────────┐
│  FastAPI Server (Port 8000)              │
│  - WebSocket for real-time updates      │
│  - Multi-instance management (1-9)      │
│  - RESTful API                           │
└──────────────┬──────────────────────────┘
               │
┌──────────────▼──────────────────────────┐
│  AI Agent Core                           │
│  - LLM Integration (Gemini/LM Studio)   │
│  - Task Planning & Execution             │
│  - Memory Management                     │
└──────────────┬──────────────────────────┘
               │
┌──────────────▼──────────────────────────┐
│  Browser Manager (Playwright)            │
│  - Headless Chromium                     │
│  - Screenshot Capture                    │
│  - Page Interaction                      │
└─────────────────────────────────────────┘

📚 Features

Core Capabilities

  • Autonomous Navigation: AI decides where to click, what to type
  • Multi-Step Workflows: Chain complex actions together
  • Context Awareness: Remembers previous actions and decisions
  • Screenshot Analysis: AI "sees" the page and makes decisions
  • Text Extraction: Reads page content for better understanding
  • Session Persistence: Maintains cookies and login state

Advanced Features

  • 9-Instance Grid: Run up to 9 browser agents simultaneously
  • Headless Mode: Runs in background without visible browser
  • CDP Support: Chrome DevTools Protocol for advanced control
  • Memory System: Stores user preferences and company context
  • Streaming Logs: Real-time visibility into agent actions

🛠️ Development

Project Structure

browser_agent/
├── core/
│   ├── agent.py          # Main AI agent logic
│   ├── browser.py        # Playwright browser manager
│   ├── llm.py           # LLM integration
│   └── memory.py        # Context & memory management
├── server/
│   ├── main.py          # FastAPI server
│   └── static/          # Web UI
└── tools/               # Agent tools (click, type, etc.)

launcher/                # Electron desktop app
├── main.js             # Electron main process
├── index.html          # Loading screen
└── package.json        # Dependencies

API Endpoints

Start a task:

curl -X POST http://localhost:8000/api/start/1 \
  -H "Content-Type: application/json" \
  -d '{"task":"Search for AI news on Hacker News"}'

Stop an agent:

curl -X POST http://localhost:8000/api/stop/1

Health check:

curl http://localhost:8000/api/diagnostic

WebSocket Events

Connect to ws://localhost:8000/ws/1 for real-time updates:

  • status - Agent status changes
  • log - Action logs
  • screenshot - New screenshots
  • error - Error messages

🤝 Contributing

We welcome contributions! This is an open-source project maintained by the TaskCato team.

How to Contribute

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Areas We Need Help

  • 🐛 Bug fixes and stability improvements
  • 📝 Documentation and examples
  • 🔌 New tool integrations
  • 🧪 Test coverage
  • 🎨 UI/UX improvements

📖 Use Cases

Web Scraping

"Go to ProductHunt and extract the top 10 products today"

Form Automation

"Fill out the contact form on example.com with my info"

Research

"Research the pricing pages of Stripe, Square, and PayPal"

Testing

"Test the login flow on staging.myapp.com"

Data Collection

"Find email addresses of CTOs at Series A fintech startups"

⚙️ Configuration

Environment Variables

# LLM Provider
LLM_PROVIDER=lm_studio  # or 'gemini'
GEMINI_API_KEY=your_key_here

# Browser Settings
HEADLESS=true  # Run browser in headless mode
VIEWPORT_WIDTH=1280
VIEWPORT_HEIGHT=800

# Server
PORT=8000
HOST=0.0.0.0

Settings File

Settings are stored in browser_agent/data/settings.json:

{
  "provider": "lm_studio",
  "lm_studio_url": "http://localhost:1234/v1",
  "lm_studio_model": "qwen/qwen3-vl-8b",
  "gemini_api_key": "",
  "gemini_model": "gemini-2.0-flash-exp"
}

🔒 Security & Privacy

  • No Data Collection: All data stays on your machine
  • Local LLM Support: Use LM Studio for complete privacy
  • Session Isolation: Each browser instance is isolated
  • No Telemetry: We don't track usage

🐛 Known Issues

  • UI Message Display: Agent responses may not appear in chat (check terminal logs)
  • Browser View Content: Grid cells show blank instead of live browser content
  • Windows Support: Electron app tested primarily on macOS

See Issues for full list.

📄 License

MIT License - see LICENSE file for details.

🌟 About TaskCato

This browser agent is the core technology behind TaskCato.com - an AI co-founder platform that helps founders automate their entire business workflow.

TaskCato adds:

  • 15+ SaaS integrations (HubSpot, Gmail, Slack, etc.)
  • AI voice calling capabilities
  • Advanced persistent memory
  • Team collaboration features
  • Managed cloud hosting
  • Priority support

Try TaskCato free: https://taskcato.com

💬 Community

🙏 Acknowledgments

Built with:


Made with ❤️ by the TaskCato team

Building the future of autonomous work, one browser action at a time.

About

Local browser automation agent powered by LM Studio — fully local LLMs

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages