Apache Cassandra is a high-performance column-family store and can scale…

Question

0

Asked: May 15, 20262026-05-15T13:57:04+00:00 2026-05-15T13:57:04+00:00

I’m just getting started with HTMLUnit and what I’m looking to do is take

0

I’m just getting started with HTMLUnit and what I’m looking to do is take a webpage and extract out the raw text from it minus all the html markup.

Can htmlunit accomplish that? If so, how? Or is there another library I should be looking at?

for example if the page contains

<body><p>para1 test info</p><div><p>more stuff here</p></div>

I’d like it to output

para1 test info more stuff here

thanks

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-05-15T13:57:05+00:00

http://htmlunit.sourceforge.net/gettingStarted.html indicates that this is indeed possible.

@Test
public void homePage() throws Exception {
    final WebClient webClient = new WebClient();
    final HtmlPage page = webClient.getPage("http://htmlunit.sourceforge.net");
    assertEquals("HtmlUnit - Welcome to HtmlUnit", page.getTitleText());

    final String pageAsXml = page.asXml();
    assertTrue(pageAsXml.contains("<body class=\"composite\">"));

    final String pageAsText = page.asText();
    assertTrue(pageAsText.contains("Support for the HTTP and HTTPS protocols"));
}

NB: the page.asText() command seems to offer exactly what you are after.

Javadoc for asText (Inherited from DomNode to HtmlPage)

How to approach applying for a job at a company ...

How to handle personal stress caused by utterly incompetent and ...

What is a programmer’s life like?

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions