Learn VBA & Macros in 1 Week!

PHP - Parsing With Php Simple Html Dom Parser - A Heritage

Full Excel VBA Course - Beginner to Expert

Parsing With Php Simple Html Dom Parser - A Heritage	View Content

Hello dear Community,

i have a document i need to parse it and spit out only this part of the table:

see http://schulnetz.nibis.de/db/schulen/schule.php?schulnr=67003&lschb=

how to i parse the stuff!? With perl or php?

Note i have the xpaths (see below) Sad that i cannot apply them on Simple DOM Parser since this Dom Parser does not work with Xpaths but with CSS-Selectors:

Well i want to get all the data with that are within the table that name is called class="fliess"

How to dump all the results?
BTW - thinking about the most elegant way, i think it is the most pretty way would be to do it with perl - So i can try it with HTML::TableExtract or....

Well what do you suggest - Which way to choose to do this [very] simple thing?

Look forward to hear from you!

see the xpaths:

Schule: /html/body/center/table/tbody/tr[2]/td[1]
Stasse: /html/body/center/table/tbody/tr[3]/td[1]
Ort: /html/body/center/table/tbody/tr[4]/td[1]
Tel: /html/body/center/table/tbody/tr[5]/td[1]
Schulgliederungen: /html/body/center/table/tbody/tr[6]/td[1]
Besonderheite: /html/body/center/table/tbody/tr[7]/td[1]
E-Mail: /html/body/center/table/tbody/tr[8]/td[1]
Schulnummer: /html/body/center/table/tbody/tr[9]/td[1]

Full Excel VBA Course - Beginner to Expert

Php Simple Html Dom Parser Fails On A Simple Example - [driving Me Nuts]

Similar Tutorials

View Content

Hi everyone,

I'm trying to select either a class or an id using PHP Simple HTML DOM Parser with absolutely no luck. My example is very simple and seems to comply to the examples given in the manual(http://simplehtmldom.sourceforge.net/manual.htm) but it just wont work, it's driving me up the wall.

Here is my example: http://schulnetz.nibis.de/db/schulen/schule.php?schulnr=94468&lschb=

I think the HTML is invalid: i cannot parse it.

Well i need more examples - probly i have overseen something!

If anybody has a working example of Simple-html-dom-parser...i would be happy.
The examples on the developersite are not very helpful.

your dilbertone

Simple Html Parser Help

Similar Tutorials

View Content

im using simple_html_dom.php

i want to extract the following html:
and number the array key so i will know the location of each <td> and extract the value the
this cell:
<TD ALIGN=RIGHT NOWRAP class="ftableline1">
3.7200
</TD>
with this :
Code: [Select]
foreach($html->find('td[class=ftableline1]') as $e)
echo $e->innertext . '<br>';

Code: [Select]
<TR class="ftableline1">

<TD ALIGN=RIGHT NOWRAP class="ftableline1">
3.7200
</TD>
<TD ALIGN=RIGHT NOWRAP class="ftableline1">

3.5400

</TD>
<TD ALIGN=RIGHT NOWRAP class="ftableline1">
3.6651
</TD>

<TD ALIGN=RIGHT NOWRAP class="ftableline1">

3.5982

</TD>
<TD align="right" NOWRAP class="ftableline1">

<A HREF=_matbea=1><IMG SRC="images/tezuga_graphit.gif" WIDTH=15 HEIGHT=15 ALT="Show Graph" BORDER="0"></a><BR>

</TD>
<TD ALIGN=RIGHT NOWRAP>0.01%</TD>
<TD ALIGN=right dir="rtl">

<IMG SRC="images/arrow_up.gif" WIDTH=10 HEIGHT=8 BORDER=0><BR>

</TD>
<TD align="right" NOWRAP dir="rtl" class="ftableline1">

3.6316

</TD>
<TD align="right" NOWRAP dir="rtl" class="ftableline1">

1

</TD>
<TD ALIGN=RIGHT NOWRAP dir="rtl" class="ftableline1">

<A HREF=_matbea=1> דולר ארה"ב</A><BR>

</TD>

<TD align="right" NOWRAP dir="rtl" class="ftableline1">

<A href="_matbea=1"><IMG SRC="../../meida/images/f1.gif" HEIGHT=15 WIDTH=21 border=0></A><BR>

</TD>
<TD ALIGN=center NOWRAP dir="rtl"><INPUT TYPE="Checkbox" VALUE="1" NAME="check" id="check" ></TD>

</TR>

Need Help For Fixing Html When Scrape Using Php Simple Html Dom Parser

Similar Tutorials

View Content

require_once 'phpSimpleHtmlDomClass.php'; $html = '<div> <div class="man">Name: madac</div> <div class="man">Age: 18 <div class="man">Class: 12</div> </div>' $name=$html->find('div[class="man"]', 0)->innertext; $age=$html->find('div[class="man"]', 1)->innertext; $cls=$html->find('div[class="man"]', 2)->innertext;

wanna get a text from each div class="man" but it didn't work because there is a missing closing div tag on 2nd line of html code. please help me to fix this.

thanks in advance.

Php Simple Html Dom Parser And Database Insertion

Similar Tutorials

View Content

Hi All,

I am using the PHP Simple HTML DOM parser to connect to a financials website, parse out a companies financial information (Income statement in this case) and then insert the scrapped data into a mysql database that I can then later use to run automated calculations.

Here is the code I have so far:

Code: [Select]
<?php
include_once 'simple_html_dom.php';

//Connect to financial Website and Create DOM from URL
$income_statement = file_get_html('http://www.WEBSITE.com/finance?etc..etc...etc...etc...');

//PULL FINANCIAL DATA
foreach($income_statement->find('td[class]' ) as $lines=>$data) {

echo $data->plaintext . "<br/>";

}

// clean up memory
$html->clear();
unset($html);
?>

So far I am able to get output that looks like this:

Code: [Select]
Revenue
336.57
331.52
324.32
319.29
320.40
Other Revenue, Total
-
-
-
-
-
Total Revenue
336.57
331.52
324.32
319.29
320.40
etc.............................
But being a newb I do not understand how I can break each $ value and each - into their own variables and then insert them to their corresponding mysql table fields. During the database insert I would like to ignore field headings from insertion (i.e Revenue, Total Revenue, etc....

Any help would be absolutely amazing, as I have been reading, scripting and searching for information like crazy, but just can't seem to figure it out.

Simple Html Dom Parser 'simple_html_dom.php' Problem

Similar Tutorials

View Content

Php Simple Html Dom Parser - How To Get Up To Speed With This Approach?

Similar Tutorials

View Content

hello dear community,

i am currently wroking on a approach to parse some sites that contain datas on Foundations in Switzerland
with some details like goals, contact-E-Mail and the like,,,

See http://www.foundationfinder.ch/ which has a dataset of 790 foundations. All the data are free to use - with no limitations copyrights on it.

I have tried it with PHP Simple HTML DOM Parser - but , i have seen that it is difficult to get all necessary data -that is needed to get it up and running.

Who is wanting to jump in and help in creating this scraper/parser. I love to hear from you.

Please help me - to get up to speed with this approach?

regards
Dilbertone

Simple Html Dom Parser: Starting Points For A Very Easy Example

Similar Tutorials

View Content

Hello dear friends,

first of all : merry merry Xmas!!!

i want to parse with the simple Simple HTML DOM Parser,

well i am pretty new to php and to the Simple HTML DOM Parser.

My example: http://schulen.bildung-rp.de/gehezu/startseite/einzelanzeige.html?tx_wfqbe_pi1[uid]=60119

I want to collect the data in the block:

I have investigated the sourcecode - and found out that the attribute of interest should be this one: class="content"div class="content">

here the code is: - my trails.

// inculde the Simple HTML DOM Parser
include_once('simple_html_dom.php');

// get the file we want to parse right now,create a DOM
$html = file_get_html('');

// simple_html_dom::find() creates a new
// simple_html_dom-Objekt, that consists out of
// corresponding childelements

foreach($html->find('class: content ') as $h3) {

  // simple_html_dom::get the text in a tag
  // den Text innerhalb eines Tags
  if($h3->innertext == 'Text of a H3 Tag') {
    break;
  }
}

// simple_html_dom::next_sibling() gives the
// next   Element
$table = $h3->next_sibling();

but believe me - it gives me not back what is aimed.

what have id done wrong...?

dbone

Native Php Dom Extension Versus Simple Dom Html Parser

Similar Tutorials

View Content

good day dear community,

this is a big issue. I have to decide: between native PHP DOM Extension or of simple DOM html parser

well i want to parse the site he http://buergerstiftungen.de/cps/rde/xchg/SID-A7DCD0D1-702CE0FA/buergerstiftungen/hs.xsl/db.htm

http://buergerstiftungen.de/cps/rde/xchg/SID-A7DCD0D1-702CE0FA/buergerstiftungen/hs.xsl/db.htm

I will suggest to use the native PHP "DOM" Extension instead of "simple html parser", since it will be much faster and easier

What do you think about this one here...:

Code: [Select]
$doc = new DOMDocument
@$doc->loadHTMLFile('...URL....'); // Using the @ operator to hide parse errors
$contents = $doc->getElementById('content')->nodeValue; // Text contents of #content

look forward to hear from you

best regards
db1

Php Simple Html Dom Parser - Compile Error In Php 5.2 When Used As Object

Similar Tutorials

View Content

I'm using PHP 5.2 Server and Simple HTML DOM 1.5. This script scrape or extract data from a football site, its fully working on PHP 5.9 Server but I need to know how I can fix it for PHP 5.2 server. Can someone give me a hint on how can I fix the error? Thanks in advance.

My PHP 5.2 Server script output shows:
++++++++++++++++
Object id #599 Object id #604 Object id #609 Object id #614 Object id #619
Object id #627 Object id #632 Object id #637 Object id #642 Object id #647
Object id #655 Object id #660 Object id #665 Object id #670 Object id #675
Object id #683 Object id #688 Object id #693 Object id #698 Object id #703
Object id #711 Object id #716 Object id #721 Object id #726 Object id #731
++++++++++++++++

while PHP 5.9 Server says
++++++++++++++++
Rk Player Team POS OPPONENT
1 Aaron Rodgers GB QB at CAR
2 Tom Brady NE QB vs. SD
3 Matt Schaub HOU QB at MIA
4 Michael Vick PHI QB at ATL
++++++++++++++++

I did applied the bug solution listed on https://sourceforge.net/tracker/index.php?func=detail&aid=3107230&group_id=218559&atid=1044037 but it is still not working. It says:
++++++++++++++++
Details:

I get compiler errors in PHP 5.2 when using this as an object.

The offending lines are 609 and 940, which both contain this construct:

if ($this->size>0) $this->char = $this->doc[0];

This tries to get the first character of $this->doc, but PHP 5.2 sees it as trying to access it as an array. It's easily fixed by this:

if ($this->size>0) $this->char = substr($this->doc, 0, 1);

Or you could probably use chr(ord($this->doc)) as well. Either way solves the compile error without changing functionality.
++++++++++++++++

Here are my codes:

Code: [Select]
<?php
# don't forget the library
include('simple_html_dom.php');

# this is the global array we fill with article information
$articles = array();
$source = 'http://www.athlonsports.com/columns/winning-game-plan/fantasy-football-qb-rankings';
# passing in the first page to parse, it will crawl to the end
# on its own
getArticles($source);

function getArticles($page) {
global $articles, $descriptions;

$html = new simple_html_dom();
$html->load_file($page);

//$items = $html->find('div[class=preview]');
$items = $html->find('tbody tr');

foreach($items as $post) {
    # remember comments count as nodes
   /*$articles[] = array($post->children(3)->outertext,
                        $post->children(6)->first_child()->outertext);*/
    $articles[] = array($post->children(0), $post->children(1), $post->children(2), $post->children(3), $post->children(4));
}

# lets see if there's a next page
if($next = $html->find('a[class=nextpostslink]', 0)) {
    $URL = $next->href;
    echo "going on to $URL <<<\n";
    # memory leak clean up
   $html->clear();
    unset($html);

    getArticles($URL);
}
}

?>

<html>
<head>
</head>
<body>
<?
echo "Source: " . $source;
?>
<table cellpadding="5" cellspacing="0" border="0">
<?php
    foreach($articles as $item) {
        echo "<tr>";
        echo "<td>" . $item[0] . "</td><td>" . $item[1] . "</td><td>" . $item[2] . "</td>";
        echo "<td>" . $item[3] . "</td><td>" . $item[4] . "</td>";
        echo "<tr>";
    }
?>
</table>

</body>
</html>

Php Html Dom Parser

Similar Tutorials

View Content

Im using some software called php html dom parser i wont to be able to keep the souce tidy

i.e before dom parser

<?php

//////////////////////SEO TOOL///////////////////////////

$title = 'Green Deal Nationwide - PB Energy Solutions Ltd';
$description = 'Delivering all your environmental needs to \'green\' up your business, improve reputation, increase profitability and give a competitive advantage.';

///////////////////////////////////////////////////////

?>
<?php include('includes/settings.php'); ?>
<?php include('includes/header.php'); ?>

<div class="container">

<div id="large-page-img">
<img src="<?php echo URL(); ?>images/home-page-slide.jpg" width="911" height="230" />
<img src="<?php echo URL(); ?>images/home-page-slide-1.jpg" width="911" height="230" />
<img src="<?php echo URL(); ?>images/home-page-slide-2.jpg" width="911" height="230" />
</div>
<div id="content-home">

<div class="iedit">

after dom parser saved to file

<?php //////////////////////SEO TOOL/////////////////////////// $title = 'Green Deal Nationwide - PB Energy Solutions Ltd'; $description = 'Delivering all your environmental needs to \'green\' up your business, improve reputation, increase profitability and give a competitive advantage.'; /////////////////////////////////////////////////////// ?> <?php include('includes/settings.php'); ?> <?php include('includes/header.php'); ?> <div class="container"> <div id="large-page-img"> <img src="<?php echo URL(); ?>images/home-page-slide.jpg" width="911" height="230" /> <img src="<?php echo URL(); ?>images/home-page-slide-1.jpg" width="911" height="230" /> <img src="<?php echo URL(); ?>images/home-page-slide-2.jpg" width="911" height="230" /> </div> <div id="content-home"> <div class="iedit"><div class="iedit">

is there anyway i can keep it like the original fil after dom?

Html Dom Parser

Similar Tutorials

View Content

Parsing Xml With Simple Xml

Similar Tutorials

View Content

I'm trying to parse an XML file with Simple XML

Here's a part of the XML:

Code: [Select]
<CueList xmlns="urn:CueListSchema.xml" xmlns:s="urn:schemas-rcsworks-com:SongSchema" xmlns:n="urn:schemas-rcsworks-com:NoteSchema" xmlns:l="urn:schemas-rcsworks-com:LinkSchema" xmlns:t="urn:schemas-rcsworks-com:TrafficSchema" xmlns:p="urn:schemas-rcsworks-com:ProductSchema" xmlns:m="urn:schemas-rcsworks-com:MediaSchema" xmlns:w="urn:schemas-rcsworks-com:WebPageSchema" xmlns:ns="urn:CueListSchema.xml" time="2011-11-02T14:56:35">
<Event eventID="2" eventType="song" status="happening" scheduledTime="14:56:31" scheduledDuration="217.00">
<s:Song title="Mr.Brightside" internalID="0077000000025FA50000">
<s:Artist name="The Killers" sequenceNumber="1" internalID="0067000000020CB20000" sortName="The Killers"/>
<m:Media ID="{A9E1341C-A638-4048-B8B5-199667E69FA5}" runTime="217.03" fileName="{A9E1341C-A638-4048-B8B5-199667E69FA5}.wav"/>
</s:Song>
</Event>
</CueList>
Now, I've got it to read the first part <event>

foreach ($CueList->Event as $event) {$type = $event['eventType'];}()

but I now want to read the Song Title on the next row down, and I'm having an issue with the s:Song part?

Any idea how I would do this, or even what you'd call that so I can look it up?

Thanks,
David

Create Html Parser Loop Through

Similar Tutorials

View Content

how should i approach the following:
a page with a products list+link to product page

i want to build a crawler that loops through all the products in the list and goes to the product page and
and parses the product page.

need help with the loop

Portiing Over A Parser From Bs4 To Simplehtmldom-parser

Similar Tutorials

View Content

hello dear Freaks

i am currently musing bout the portover of a python bs4 parser to php - working with the simplehtmldom-parser / pr the DOM-selectors... (see below).

The project: for a list of meta-data of wordpress-plugins: - approx 50 plugins are of interest! but the challenge is: i want to fetch meta-data of all the existing plugins. What i subsequently want to filter out after the fetch is - those plugins that have the newest timestamp - that are updated (most) recently. It is all aobut acutality...

https://wordpress.org/plugins/participants-database ....and so on and so forth.

https://wordpress.org/plugins/wp-job-manager
https://wordpress.org/plugins/ninja-forms
https://wordpress.org/plugins/participants-database ....and so on and so forth.

we have the following set of meta-data for each wordpress-plugin:

Version: 1.9.5.12 
installations: 10,000+    
WordPress Version: 5.0 or higher 
Tested up to: 5.4 PHP  
Version: 5.6 or higher    
Tags 3 Tags:databasemembersign-up formvolunteer
Last updated: 19 hours ago

the project consits of two parts: the looping-part: (which seems to be pretty straightforward). the parser-part: where i have some issues - see below. I'm trying to loop through an array of URLs and scrape the data below from a list of wordpress-plugins. See my loop below-

as a base i think it is good starting point to work from the following target-url:

plugins wordpress.org/plugins/browse/popular with 99 pages of content: cf ...
wordpress.org/plugins/browse/popular/page/1
wordpress.org/plugins/browse/popular/page/2
wordpress.org/plugins/browse/popular/page/99

the Output of text_nodes:

['Version: 1.9.5.12', 'Active installations: 10,000+', 'Tested up to: 5.6 ']

but if we want to fetch the data of all the wordpress-plugins and subesquently sort them to show the -let us say - latest 50 updated plugins. This would be a interesting task:

first of all we need to fetch the urls

then we fetch the information and have to sort out the newest- the newest timestamp. Ie the plugin that updated most recently

List the 50 newest items - that are the 50 plugins that are updated recently ..

we have the following set

see here the Soup_

 soup = BeautifulSoup(r.content, 'html.parser')
        target = [item.get_text(strip=True, separator=" ") for item in soup.find(
            "h3", class_="screen-reader-text").find_next("ul").findAll("li")[:8]]
        head = [soup.find("h1", class_="plugin-title").text]
        new = [x for x in target if x.startswith(
            ("V", "Las", "Ac", "W", "T", "P"))]
        return head + new


with ThreadPoolExecutor(max_workers=50) as executor1:
    futures1 = [executor1.submit(parser, url) for url in allin]

for future in futures1:
    print(future.result())

see the formal output

Quote

[lorem ipsum dolor sit amet', 'Version: 2.34.1', 'Last updated: 5 months ago', 'Tags: magna aliquyam erat, sed diam voluptua. At vero eos et accusam']
[consetetur sadipscing elitr', 'Version: 6.54.1', 'Last updated: 5 months ago', 'Tags: lorem ipsum dolor sit amet']
[sed diam nonumy eirmod tempor invidunt ut labore', 'Version: 7.16.1', 'Last updated: 5 months ago', 'Tags: tarifa, sevilla lisabin invidunt ut labore et dolore magna aliquyam erat']
[tempor invidunt ut taria malaga jerusalem labore', 'Version: 9.58.1', 'Last updated: 5 months ago', 'Tags: ilabore et lissabon dolore magna aliquyam erat']

background: https://stackoverflow.com/questions/61106309/fetching-multiple-urls-with-beautifulsoup-gathering-meta-data-in-wp-plugins

Well - i guess that we c an do this with the simple DOM Parser - here the seclector reference.

https://stackoverflow.com/questions/1390568/how-can-i-match-on-an-attribute-that-contains-a-certain-string

look forward to any hint and help.

have a great day

Edited May 3, 2020 by dil_bert

Parsing Html Requests

Similar Tutorials

View Content

Hi,

I am trying to make a web interface for a robot, I have written php to send/recieve values via a serial port to my robot. They work.

I am now tring to develop my web interface.

I'm using java to generate http requests client side in the form of;
Code: [Select]
/request?command=Forward&param1=254
I was wondering how I can parse the command and param1 in php sereverside?

Or is there a better alternative?

Parsing Html From Wikipedia

Similar Tutorials

View Content

Hi guys. I have been using the wikipedia API to retrieve information about a topic. Ive managed to get a response and retrieve the first section of the topic (in this case football)

Using this method - http://en.wikipedia.org/w/api.php?action=parse&page='.$search.'&redirects=1&format=json&prop=text&section=0');

However the first section that is retrieved includes the pictures and i just want to main text from the introduction.

The code that is sent back from wiki is this -
Code: [Select]
Array
(
[parse] => Array
(
[text] => Array
(
[*] => <div class="dablink">This article is about sports known as football. For the ball used in these sports, see <a href="/wiki/Football_(ball)">Football (ball)</a>.</div>
<div class="thumb tright">
<div class="thumbinner" style="width:227px;"><a href="/wiki/File:Football4.png" class="image"><img alt="" src="http://upload.wikimedia.org/wikipedia/commons/thumb/d/d2/Football4.png/225px-Football4.png" width="225" height="274" class="thumbimage" /></a>
<div class="thumbcaption">
<div class="magnify"><a href="/wiki/File:Football4.png" class="internal" title="Enlarge"><img src="http://bits.wikimedia.org/skins-1.17/common/images/magnify-clip.png" width="15" height="11" alt="" /></a></div>
Some of the many different games known as football. From top left to bottom right: <a href="/wiki/Association_football">Association football</a> or soccer, <a href="/wiki/Australian_rules_football">Australian rules football</a>, <a href="/wiki/International_rules_football">International rules football</a>, <a href="/wiki/Rugby_Union" class="mw-redirect" title="Rugby Union">Rugby Union</a>, <a href="/wiki/Rugby_League" class="mw-redirect" title="Rugby League">Rugby League</a>, and <a href="/wiki/American_Football" class="mw-redirect" title="American Football">American Football</a>.</div>
</div>
</div>
<p>The game of <b>football</b> is any of several similar <a href="/wiki/Team_sport" title="Team sport">team sports</a>, of similar origins which involve advancing a ball into a goal area in an attempt to score. Many of these involve <a href="/wiki/Kick_(football)" title="Kick (football)">kicking</a> a ball with the foot to score a <a href="/wiki/Goal_(sport)" title="Goal (sport)">goal</a>, though not all codes of football using kicking as a primary means of advancing the ball or scoring. The most popular of these sports worldwide is <a href="/wiki/Association_football">association football</a>, more commonly known as just "football" or "soccer". Unqualified, the word <i><a href="/wiki/Football_(word)" title="Football (word)">football</a></i> applies to whichever form of football is the most popular in the regional context in which the word appears, including <a href="/wiki/American_football">American football</a>, <a href="/wiki/Australian_rules_football">Australian rules football</a>, <a href="/wiki/Canadian_football">Canadian football</a>, <a href="/wiki/Gaelic_football">Gaelic football</a>, <a href="/wiki/Rugby_league">rugby league</a>, <a href="/wiki/Rugby_union">rugby union</a> and other related games. These variations are known as "codes".</p>

I want the code that resides in the <p> tags. How would i go about parsing this and removing the rest. ive tried to get to work simple html dom parser but with no luck.

Any help would be greatly appreciated

Thanks,

DIM3NSION

Php Parsing Html Table

Similar Tutorials

View Content

Hi guys,
im trying to parse a html table from an existing website to my own. However ive run into a few problems. Does anyone know how to parse html tables?? im using the PHP DOM Parser but at the moment i am only able to return all the data on the website rather then the specific table.
Thanks for any help!

Parsing An Entire Html Table

Similar Tutorials

View Content

Hello again,

I'm trying to scrape a table from another website using preg_match, especifically, using this code:
Code: [Select]
<?php
$data = file_get_contents('http://tvcountdown.com/index.php');
$regex = '/[color=red]<table class="episode_list_table">[/color] (.+?) [color=red]</table>[/color]/';
preg_match($regex,$data,$match);
var_dump($match);
echo $match[0];
?>
Heres the thing. It doesnt work
I think it's because the first and second anchors are html tags, 'cause if I parse some other stuff without any tag, there's no problemo.

Any hints, mates?
Thanks

Moved: Regex And Html Parsing!

Similar Tutorials

View Content

This topic has been moved to PHP Regex.

http://www.phpfreaks.com/forums/index.php?topic=308636.0

Parsing Html To Strip Required Data

Similar Tutorials

View Content

So I have an interesting one for you guys this AM,
I first want to make it very clear that I am not scraping code, rather I am scraping data that is needed to import into a shopping cart system for someone.
I have a URL that I am trying to scrape required data off of, however it is not returning all the data that I want. I have created a function that uses preg_match_all() and regex and I am still having issues striping what I want.

here is a link to my test what I am wanting to strip from http://visualrealityink.com/dev/clients/rug_src/scrapeing/Rugsource/www.vendio.com/stores/Rugsource1/item/other/tribal-wool-3x5-shiraz-persian/lid=10363581.html

I am wanting to grab all this data:

Quote

Item Number:
K-686
Style :    Shiraz
Province :    Fars
Made In :    Iran
Foundation :    Wool
Pile :    100% Wool
Colors :    Red, Navy Blue, Ivory, Forest Green, Light Blue, Orange
Size (feet) :    4' 11" x 3' 4"
Size (Centimeter) :    155 x 103
Age :    20-25 Years Old
Condition :    Very Good
KPSI (knots per sq. inch) :    130 knots per square inch
Woven :    Hand Knotted
Shipping and Handling :    Free Shipping(For Mainland USA)
Est. Retail Value :    $2,700.00

Here is the code note that $url holds the link above.
Code: [Select]
$html = file_get_contents($url);

$newlines = array("\t","\n","\r","\x20\x20","\0","\x0B");

$html = str_replace($newlinews, "", html_entity_decode($html));
preg_match_all('/<tr><td width="50%" align="right"><font color="#800000"><b>[^\s ](.*?)<\/b><\/font><\/td><td width="50%" align="left">[^\s ](.*?)<\/td><\/tr>/', $html, $matches, PREG_SET_ORDER);

foreach($matches_label as $match){
$count = 0;
echo $match[$count];
echo "<br>";
$count++;

}
echo $count;
This returns the following
Quote

Style : Shiraz
Province : Fars
Foundation : Wool
Colors : Red, Navy Blue, Ivory, Forest Green, Light Blue, Orange
Size (feet) : 4' 11" x 3' 4"
Size (Centimeter) : 155 x 103
Age : 20-25 Years Old
Condition : Very Good
Est. Retail Value : $2,700.00
1

it is missing: Quote

Inventory Number : xxxxxxx
Made In: xxxxxxxx
Pile : xxxxxxxxxx
KPSI(Knots Per Inch) : xxxxxxxxxx
Woven : xxxxxxxxx
Shopping : xxxxxxxxxxx

You can see the script in action here -> http://visualrealityink.com/dev/clients/rug_src/scrapeing/scrape_tst.php

Thanks in advance for all of your help