A PROJECT OF AVF TEAM 1
What our crawler is, what it reads, how to contact us, and how to opt out.
EOB-CatalogueBot sends this user agent on every request, without exception:
EOB-CatalogueBot/1.0 (+https://employmentoffbase.org/crawler/catalogue-bot; catalogue@employmentoffbase.org)
We never rotate it. We never send a browser user agent. We never disguise the crawler as a person. If a request to your site does not carry that string, it is not us.
AVF Team 1 is a nonprofit. Employment Off Base is its technical arm and operates this crawler on its behalf.
The work is a research and workforce project for separating service members and military spouses, who move often and lose career continuity each time.
Non-credit workforce programs are not collected by any federal dataset. IPEDS covers credit completions. A short welding certificate or a CDL course that would actually help someone arriving in your area is invisible to the federal record. We are assembling that inventory and it is free to use.
Public course listing and course detail pages, and only these fields:
That is the whole list.
We do not create accounts. We do not fill in forms. We do not execute registrations.
We obey robots.txt. If your robots.txt disallows a path, we do not fetch it. We do not fetch it from another address either.
We take the slower of the two pacing rules. Our own floor is one request per host every 2 seconds. If your robots.txt publishes a Crawl-delay, we use yours whenever yours is slower. CourseStorm publishes 10 seconds and we use 10. ed2go publishes 20 and we use 20.
We stop at a block. If we receive a 403, a 429, or a robots disallow, we record it, we stop, and a person writes to you. We do not rotate addresses, change identity, retry from elsewhere, or route around a block in any form. A block is an answer and we treat it as one.
We re-read rarely. A catalogue is re-read no more than once every 90 days. Course inventories do not change hourly and we do not pretend they do. You can slow us further or stop us entirely with one line in robots.txt.
We identify one crawler. There is no second bot, no residential proxy pool, no third-party scraping vendor acting for us.
We publish the fields listed above alongside the name of your institution and a link to your page. We cite you as the source every time.
We do not determine eligibility for any funding program. We surface options and cite the institution as the source. Whether a program qualifies for Workforce Pell, WIOA, MyCAA, GI Bill benefits or anything else is determined by the institution, the state and the federal government, and never by us.
If we cannot confirm a fact, we record it as unresolved. We do not fill a gap with a guess.
Any one of these works, and none of them requires contacting us first.
1. robots.txt. Add:
User-agent: EOB-CatalogueBot Disallow: /
We honor it on the next pass and stop.
2. Email. Write to catalogue@employmentoffbase.org and say what you want removed. We remove it and confirm. You do not need to give a reason.
3. Narrow it instead of blocking it. If some pages are fine and others are not, name the paths and we will restrict ourselves to what you allow.
Opting out costs you nothing and we do not follow up to change your mind.
Send us your catalogue and we will not crawl you at all.
A CSV, a spreadsheet, an existing feed, or an API you already publish. Roughly an hour of developer time once, and you keep control of exactly what is published about your programs. It is more accurate than anything we could read off a page, and it is the route we prefer.
Write to catalogue@employmentoffbase.org.
catalogue@employmentoffbase.org
A person reads these and replies. Not a ticket queue, not an autoresponder.
If EOB-CatalogueBot has caused a problem on your servers, tell us and we will slow down or stop the same day.