A fast and efficient tool for extracting categories and articles from any public Front knowledge base. It helps teams collect structured help center content for analysis, migration, or documentation workflows with minimal setup.
Created by Bitbash, built to showcase our approach to Scraping and Automation!
If you are looking for front-knowledge-base you've just found your team — Let’s Chat. 👆👆
This project extracts structured content from public Front knowledge bases, including categories and articles, in a clean and predictable format. It solves the challenge of manually copying or auditing help center content by automating data collection. The tool is ideal for product teams, support managers, and developers working with Front-based documentation.
- Works with any publicly accessible Front knowledge base
- Supports category-level and article-level data collection
- Designed for speed and low operational cost
- Simple configuration with flexible crawling limits
| Feature | Description |
|---|---|
| Category Listing | Retrieves all top-level knowledge base categories. |
| Article Listing | Extracts article metadata and content from categories. |
| Full Article Export | Collects all available articles in a single run. |
| Configurable Crawling | Control limits and behavior through crawler options. |
| Proxy Compatibility | Designed to work reliably with proxy-enabled requests. |
| Field Name | Field Description |
|---|---|
| category_id | Unique identifier of the knowledge base category. |
| category_name | Display name of the category. |
| article_id | Unique identifier of the article. |
| article_title | Title of the help center article. |
| article_url | Public URL of the article. |
| article_content | Main textual content of the article. |
| last_updated | Last modification timestamp of the article. |
Front Knowledge Base/
├── src/
│ ├── index.js
│ ├── crawler/
│ │ ├── categoryCrawler.js
│ │ └── articleCrawler.js
│ ├── parsers/
│ │ ├── categoryParser.js
│ │ └── articleParser.js
│ └── config/
│ └── crawler.config.json
├── data/
│ ├── sample-input.json
│ └── sample-output.json
├── package.json
└── README.md
- Support teams use it to audit help center content, so they can identify outdated or missing articles.
- Product managers use it to export documentation, so they can migrate knowledge bases between platforms.
- Developers use it to index articles, so they can integrate help content into internal tools.
- Content strategists use it to analyze article coverage, so they can improve documentation quality.
Does this tool support sub-categories? Currently, only top-level categories are supported. Articles must exist directly under their parent category to be extracted correctly.
What happens if the article limit is too low? If the maximum request limit is lower than the total number of articles, the extraction may return no results. Setting the limit to unlimited avoids this issue.
How can I verify the knowledge base URL is valid? If the knowledge base does not exist or is not public, requests may fail due to unsupported content types.
Primary Metric: Processes hundreds of articles per minute on medium-sized knowledge bases.
Reliability Metric: Maintains a high success rate when crawler limits are correctly configured.
Efficiency Metric: Optimized request handling minimizes redundant page loads.
Quality Metric: Extracted articles retain full textual content with consistent field structure.
