Showing posts with label web crawlers. Show all posts
Showing posts with label web crawlers. Show all posts

Monday, December 31, 2012

Web scraping with node.js

Nice post on Web scraping, using node.js - but the techniques are pretty universal and really worth a read for any platform.

Thursday, December 13, 2012

You know where I'm going with this blog, right?

I want to do textual analysis on both the blog and everything it links, then search for similar things so I can just let the blog run itself, more or less.

Instead of letting HackerNews find things for me, I can do it myself and maybe start feeding HackerNews instead of the other way around.

Wednesday, December 12, 2012

PhantomJS

A full headless browser in JavaScript.  Good for testing and scraping, apparently - worth investigating.

Amazon random shopper bot

Cute.

Saturday, November 3, 2012

Grepsr

Grepsr appears to be a consulting service slash platform for repeating web scraping.  Pricing appears high.

Friday, June 15, 2012

Testing 3 million URLs

Wow.  Here's a nice article on lessons learned from policing dead links on StackExchange.