I can get all those url’s whose content/type is text/html, but If I want

Question

0

Asked: May 23, 20262026-05-23T17:11:36+00:00 2026-05-23T17:11:36+00:00

I can get all those url’s whose content/type is text/html, but If I want

0

I can get all those url’s whose content/type is text/html, but If I want those urls whose content/type is not text/html. Then how can we check that. As for the string we can use contains method, but it doesn’t have anything like notcontains.. Any suggestions will be appreciated.. And also

The key variable contains:

Content-Type=[text/html; charset=ISO-8859-1]

This is the below code to check for text/html and I tried also for content-type that are not text/html but it also prints out those whose content-type are also text/html.

    try {
            URL url1 = new URL(url);
            System.out.println("URL:- " +url1);
            URLConnection connection = url1.openConnection();

            Map responseMap = connection.getHeaderFields();
            Iterator iterator = responseMap.entrySet().iterator();
            while (iterator.hasNext())
            {
                String key = iterator.next().toString();

                if (key.contains("text/html") || key.contains("text/xhtml"))
                {
                    System.out.println(key);
                    // Content-Type=[text/html; charset=ISO-8859-1]
                    if (filters.matcher(key) != null){
                        System.out.println(url1);
                        try {
                            final File parentDir = new File("crawl_html");
                            parentDir.mkdir();
                            final String hash = MD5Util.md5Hex(url1.toString());
                            final String fileName = hash + ".txt";
                            final File file = new File(parentDir, fileName);
                            boolean success =file.createNewFile(); // Creates file crawl_html/abc.txt


                             System.out.println("hash:-"  + hash);

                                    System.out.println(file);
                            // Create file if it does not exist



                                // File did not exist and was created
                                FileOutputStream fos = new FileOutputStream(file, true);

                                PrintWriter out = new PrintWriter(fos);

                                // Also could be written as follows on one line
                                // Printwriter out = new PrintWriter(new FileWriter(args[0]));

                                            // Write text to file
                                Tika t = new Tika();
                                String content= t.parseToString(new URL(url1.toString()));


                                out.println("===============================================================");
                                out.println(url1);
                                out.println(key);
                                out.println(success);
                                out.println(content);

                                out.println("===============================================================");
                                out.close();
                                fos.flush();
                                fos.close();



                        } catch (FileNotFoundException e) {
                            // TODO Auto-generated catch block
                            e.printStackTrace();
                        } catch (IOException e) {
                            // TODO Auto-generated catch block

                            e.printStackTrace();
                        } catch (TikaException e) {
                            // TODO Auto-generated catch block
                            e.printStackTrace();
                        }


                        // http://google.com
                    }
                }
  else if (!connection.getContentType().startsWith("text/html"))//print duplicate records of each url
                //else if (!key.contains("text/html"))
                {
                    if (filters.matcher(key) != null){
                     try {
                        final File parentDir = new File("crawl_media");
                        parentDir.mkdir();
                        final String hash = MD5Util.md5Hex(url1.toString());
                        final String fileName = hash + ".txt";
                        final File file = new File(parentDir, fileName);
                     // Create file if it does not exist
                        boolean success =file.createNewFile(); // Creates file crawl_html/abc.txt


                         System.out.println("hash:-"  + hash);

                         Tika t = new Tika();
                        String content_media= t.parseToString(new URL(url1.toString()));



                             // File did not exist and was created
                            FileOutputStream fos = new FileOutputStream(file, true);

                             PrintWriter out = new PrintWriter(fos);

                             // Also could be written as follows on one line
                             // Printwriter out = new PrintWriter(new FileWriter(args[0]));

                                         // Write text to file
                             out.println("===============================================================");
                             out.println(url1);
                             out.println(key);
                             out.println(success);
                             out.println(content_media);
                             //out.println("===============================================================");
                             out.close();
                             fos.flush();
                             fos.close();




                     } catch (FileNotFoundException e) {
                         // TODO Auto-generated catch block
                         e.printStackTrace();
                     } catch (IOException e) {
                         // TODO Auto-generated catch block

                         e.printStackTrace();
                     } catch (TikaException e) {
                        // TODO Auto-generated catch block
                        e.printStackTrace();
                    }
                    }

                }



            }
        } catch (MalformedURLException e) {
            e.printStackTrace();
        } catch (IOException e) {
            e.printStackTrace();
        }



        System.out.println("=============");
    }   
}

One method is to check individually for each content-type like for pdf it is application/pdf

if (key.contains("application/pdf")

and in the same way for xml… But any other method other than this…

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-05-23T17:11:36+00:00

Editorial Team

2026-05-23T17:11:36+00:00Added an answer on May 23, 2026 at 5:11 pm

Would this help?

 if (!connection.getContentType.startsWith("text/html"))

0

Reply
Share
Share

- Report

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions

I can get all those url’s whose content/type is text/html, but If I want

Leave an answerCancel reply

1 Answer

Leave an answer
Cancel reply