1

我爬取了电影列表并将它们存储在我的数据库中。对于仅包含英文字符的电影,一切正常,但问题是某些包含非英文字符的电影名称无法正确显示。例如,意大利电影“Il più roughle dei giorni”存储为“Il piùrawle dei giorni”。

有人可以让我知道是否有任何解决方案吗?(我知道我可以设置爬虫的语言,我已经爬过意大利语的电影标题,但是当我想爬英文标题时,Imdb中还有一些电影没有英文字符)

编辑:这是我的代码:

String baseUrl = "http://www.imdb.com/search/title?at=0&count=250&sort=num_votes,desc&start="+start+"&title_type=feature&view=simple";

label1:  try {

     org.jsoup.Connection con = Jsoup.connect(baseUrl).userAgent("Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/535.21 (KHTML, like Gecko) Chrome/19.0.1042.0 Safari/535.21").header("Accept-Language", "en");
     con.timeout(30000).ignoreHttpErrors(true).followRedirects(true);
     Response resp = con.execute();
     Document doc = null;

     if (resp.statusCode() == 200) {

         doc = con.get();                                       

         Elements myElements = doc.getElementsByClass("results").first().getElementsByTag("table");
         Elements trs = myElements.select(":not(thead) tr");

         for (int i = 0; i < trs.size(); i++) {

             Element tr = trs.get(i);
             Elements tds = tr.select("td");

             for (int j = 3; j < tds.size(); j++) {

                 Elements links = tds.select("a[href]");
                 String titleId = links.attr("href");
                 String movietitle = links.html();    

                  //I ADDED YOUR CODE HERE
                   Charset c = Charset.forName("UTF-16BE");

                        ByteBuffer b = c.encode(movietitle);
                        for (int m = 0; b.hasRemaining(); m++) {
                            int charValue = (b.get()) & 0xff;
                            System.out.print((char) charValue);
                        }   

               // try{    

                //   String query = "INSERT into test (movieName,ImdbId)" + "VALUES (?,?)";
    //               PreparedStatement preparedStmt = conn.prepareStatement(query);
    //               preparedStmt.setString (1, movietitle);
      //               preparedStmt.setString (2, titleId );
       //          }catch (Exception e)
        //       {
        //           e.printStackTrace();
        //       }

谢谢,

4

1 回答 1

1

在这里,我复制粘贴了问题中共享的字符串并尝试了

public class Test {
    public static void main (String...a) throws Exception {
        String s = "Il più crudele dei giorni";
        Charset c = Charset.forName("UTF-16BE");

        ByteBuffer b = c.encode(s);
        for (int i = 0; b.hasRemaining(); i++) {
            int charValue = (b.get()) & 0xff;
            System.out.print((char) charValue);
        }
    }
}

s这会在控制台上按原样打印。我假设您已经有部分代码写入文件。如果它适合您,您可以尝试集成上述代码。

于 2014-10-05T13:09:11.707 回答