I am trying to use BeautifulSoup to get text from web pages. Below is

Question

0

Asked: June 14, 20262026-06-14T09:08:55+00:00 2026-06-14T09:08:55+00:00

I am trying to use BeautifulSoup to get text from web pages. Below is

0

I am trying to use BeautifulSoup to get text from web pages.

Below is a script I’ve written to do so. It takes two arguments, first is the input HTML or XML file, the second output file.

import sys
from bs4 import BeautifulSoup

def stripTags(s): return BeautifulSoup(s).get_text()

def stripTagsFromFile(inFile, outFile):
    open(outFile, 'w').write(stripTags(open(inFile).read()).encode("utf-8"))

def main(argv):
    if len(sys.argv) <> 3:
        print 'Usage:\t\t', sys.argv[0], 'input.html output.txt'
        return 1
    stripTagsFromFile(sys.argv[1], sys.argv[2])
    return 0

if __name__ == "__main__":
    sys.exit(main(sys.argv))

Unfortunately, for many web pages, for example: http://www.greatjobsinteaching.co.uk/career/134112/Education-Manager-Location
I get something like this (I’m showing only few first lines):

html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd"
    Education Manager  Job In London With  Caleeda | Great Jobs In Teaching

var _gaq = _gaq || [];
_gaq.push(['_setAccount', 'UA-15255540-21']);
_gaq.push(['_trackPageview']);
_gaq.push(['_trackPageLoadTime']);

Is there anything wrong with my script? I was trying to pass ‘xml’ as the second argument to BeautifulSoup’s constructor, as well as ‘html5lib’ and ‘lxml’, but it doesn’t help.
Is there an alternative to BeautifulSoup which would work better for this task? All I want is to extract the text which would be rendered in a browser for this web page.

Any help will be much appreciated.

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-06-14T09:08:56+00:00

nltk’s clean_html() is quite good at this!

Assuming that your already have your html stored in a variable html like

html = urllib.urlopen(address).read()

then just use

import nltk
clean_text = nltk.clean_html(html)

UPDATE

Support for clean_html and clean_url will be dropped for future versions of nltk. Please use BeautifulSoup for now…it’s very unfortunate.

An example on how to achieve this is on this page:

BeatifulSoup4 get_text still has javascript

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions

I am trying to use BeautifulSoup to get text from web pages. Below is

Leave an answerCancel reply

1 Answer

Leave an answer
Cancel reply