| 1 | 24,601 | 81,756 | generalscraper | Scrapes Google |
| 2 | 26,585 | 35,247 | linkedindata | Scrapes all LinkedIn profiles including terms you specify. |
| 3 | 28,288 | 81,756 | jsontochart | Take JSON files and outputs html for various types of charts |
| 4 | 30,816 | 35,247 | entityextractor | Extracts entities and terms from any JSON. |
| 5 | 31,485 | 44,980 | linkedincrawler | Crawls public LinkedIn profiles via Google |
| 6 | 34,164 | 39,155 | dircrawl | Run block on all files in dir |
| 7 | 41,295 | 81,756 | wordcloud | Takes input and outputs the same text with word size changed based on frequency. |
| 8 | 42,262 | 81,756 | uploadconvert | Converts documents to the appropriate format for Transparency Toolkit. |
| 9 | 43,552 | 81,756 | linkedinparser | Parses public LinkedIn profiles |
| 10 | 44,870 | 81,756 | parsefile | OCR file and extract metadata using Apache Tika and Tesseract |
| 11 | 49,288 | 81,756 | urlarchiver | Saves html and pdfs of websites. |
| 12 | 50,410 | 56,018 | twittercrawler | Crawls Twitter |
| 13 | 53,231 | 81,756 | extractpatterns | Extracts entities and terms from any JSON. |
| 14 | 55,332 | 56,018 | timelinegen | TimelineGen generates JSON files for use as TimelineJS data. |
| 15 | 57,390 | 81,756 | sunlightcongress | Access to Sunlight Foundation's congress data. |
| 16 | 58,339 | 81,756 | indeedparser | Parses Indeed resumes |
| 17 | 61,457 | 81,756 | jsontonetworkgraph | Generates node and link data from any JSON. |
| 18 | 63,796 | 56,018 | piplrequest | Gets data from Pipl |
| 19 | 66,325 | 81,756 | tsjobcrawler | Crawls job listing websites for jobs requiring security clearance. |
| 20 | 72,141 | 81,756 | requestmanager | Manages proxies, wait intervals, etc |
| 21 | 76,338 | 81,756 | countryconvert | Converts 2-char ISO country codes to 3-char. |
| 22 | 76,804 | 81,756 | effscraper | Scrapes EFF court documents then extracts the plaintext and metadata. |
| 23 | 77,051 | 44,980 | termextractor | Extracts entities and terms from any JSON. |
| 24 | 80,639 | 56,018 | jsontomap | Converts a JSON into a GeoJSON. |
| 25 | 81,225 | 81,756 | indeedcrawler | Crawls Indeed resumes |
| 26 | 82,203 | 81,756 | sunlightpartytime | Access to Sunlight Foundation's Party Time data. |
| 27 | 89,352 | 44,980 | acluscraper | Scrapes ACLU court documents then extracts the plaintext and metadata. |
| 28 | 92,224 | 81,756 | piplcollector | Gets data from Pipl for dir of files |
| 29 | 92,344 | 81,756 | jsoncrossreference | Crossreferences JSONs and returns the matches |
| 30 | 97,358 | 56,018 | wlsearchscraper | Gets a list of documents from the WikiLeaks search that match certain terms. |
| 31 | 108,261 | 81,756 | datacalc | Some data calculation/manipulation for Transparency Toolkit. |
| 32 | 113,061 | 81,756 | jsoncombiner | Input multiple JSONs, get back one with all the data |
| 33 | 114,669 | 81,756 | doc_integrity_check | Encrypts, verifies, and checks hashes of files |
| 34 | 116,000 | 81,756 | jsontochoropleth | Converts as JSON to a world choropleth map. |
| 35 | 118,808 | 81,756 | ttcalc | Calculation functions for Transparency Toolkit. |
| 36 | 119,250 | 81,756 | sigadparse | Extracts SIGADs from documents |
| 37 | 133,469 | 81,756 | harvesterreporter | Incremental result reporting for Transparency Toolkit |
| 38 | 151,438 | 81,756 | guardianscraper | Scrapes Guardian articles. |
| 39 | 153,118 | 56,018 | nametoemail | Gets a list of possible email addresses. |
| 40 | 153,164 | 81,756 | indeedscraper | Get resumes and job listings from indeed based on search terms and locations. |
| 41 | 175,615 | 81,756 | docintegritycheck | Encrypts, verifies, and checks hashes of files |