You cannot select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
 
 
 
Go to file
Yuri Baburov dada82099b Moved to lxml (based on decruft version); better encoding recognition. 13 years ago
readability Moved to lxml (based on decruft version); better encoding recognition. 13 years ago
.gitignore initial 14 years ago
README well that was quick; first fork added 14 years ago
setup.py made setup.py executable 14 years ago

README

This code is under the Apache License 2.0.  http://www.apache.org/licenses/LICENSE-2.0

This is a python port of a ruby port of arc90's readability project

http://lab.arc90.com/experiments/readability/

Given a html document, it pulls out the main body text and cleans it up.

Ruby port by starrhorne and iterationlabs
Python port by gfxmonk

This port uses BeautifulSoup for the HTML parsing. That means it can be
a little slow, but will work on Google App Engine (unlike libxml-based
libraries)


**note**: I don't currently have any plans for using or improving this
library, and it's far from perfect (slow, and almost certainly buggy).
So if you do something cool with it or have a better tool that does
the same job, please let me know and I can link to it from here.

If you're looking for alternatives, here's the list so far:
 - http://www.minvolai.com/blog/decruft-arc90s-readability-in-python/