What is GcrawlAI?
GcrawlAI is an open-source web scraping API that turns websites into clean Markdown, JSON, or screenshots for AI apps. Built by Chennai-based Gramosoft, it handles JavaScript rendering, stealth, and rotating proxies inside every request. Teams can use the managed cloud, billed per request, or self-host the MIT-licensed engine from GitHub.
Top Features:
- Clean Markdown: strips menus, ads, and footers so language models use fewer tokens.
- Stealth scraping: a three-tier stack escalates to residential proxies when sites block bots.
- Five endpoints: scrape, crawl, map, screenshot, and search cover most web data jobs.
Use Cases:
- RAG pipelines: feed chatbots and search indexes with fresh, structured content from websites.
- Site crawling: collect every page on a domain with depth limits and webhooks.
- Agent research: let AI agents read live web pages through the MCP server.
Who Can Use GcrawlAI?
- AI engineers: plug web data into LangChain or LlamaIndex through Python or Node.js SDKs.
- Data teams: gather product, pricing, and SEO data from many websites quickly.
- APAC companies: keep data processing in the region to meet DPDPA 2023 rules.
Pricing
- Free ($0): 50 monthly requests with JavaScript rendering, stealth, and two concurrent jobs.
- Starter ($19 per month): 3,000 monthly requests with five concurrent jobs for small projects.
- Growth ($49 per month): 50,000 monthly requests and fifteen concurrent jobs for steady workloads.
Pros and Cons
Pros:
- Simple billing: one request counts as one request, with no multipliers for rendering.
- Open source: the MIT-licensed engine can run on your own servers at no cost.
- Fair counting: failed requests are never deducted from your monthly request allowance.
Cons:
- Small free tier: fifty monthly requests cover little more than light testing and demos.
- Paid extractors: ready-made Amazon, Walmart, and Flipkart extractors sit behind paid plans.
- Young ecosystem: tutorials and third-party guides are still scarce for new users.
FAQs:
1) Is GcrawlAI open source?
Yes, the core engine is MIT-licensed on GitHub and free to self-host.
2) Does it handle JavaScript-heavy websites?
Yes, pages render in a headless browser and still count as one request.
3) What output formats are available?
Clean Markdown, raw HTML, screenshots, SEO metadata, and extracted images, chosen per request.
4) Does it work with AI agents?
Yes, Claude, GPT, and Gemini agents connect through the REST API, SDKs, or MCP.
5) Who builds the product?
Gramosoft, an ISO 27001 certified AI company based in Chennai, India.