Thursday, 10 July 2014

Large Text File Viewer

A recent issue resulted in interrogating log files > 1 Gb. All text editors I had could not cope with this including emacs, notepad++ and windows notepad. Eventually I found the simply named Large Text File Viewer from Switftgear which worked a treat. Downloadable as a zip, it requires no installation and the executable is only 572kb. Download here

Wednesday, 25 June 2014

Xpath to generate an xpath string to the current item

An xpath to be used in either xquery or XSLT to generate the heirarchical path to the current item


string-join(
   (for $node in ancestor::* 
    return 
      concat($node/name(),
              '[', 
              count($node/preceding-sibling::*[name() = $node/name()])+1, 
              ']'
            ),
      concat(name(),
              '[', 
              count(preceding-sibling::*[name() = current()/name()]) + 1, 
              ']'
            )
    ),
  '/')


Find Duplicate IDs with XSLT

This little snippet of XSLT is a useful tool to find all duplicate ids within an XML source document and generate a report with the count of duplicates and xpath to each element that has a duplciate id attribute.

This relies upon the attribute in question being names @id but it is simple enough to change this to whatever attribute you need to interrogate

Note that this is XSLT 2 and has been used with the Saxon transformation engine


<xsl:stylesheet version="2.0" 
  xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
  xmlns:xs="http://www.w3.org/2001/XMLSchema"
 
  exclude-result-prefixes="xs">

  <xsl:output indent="yes"/>
  
  <xsl:key name="ids" match="*[@id]" use="@id"/> 

  <xsl:template match="/">
    <duplicates>
      <xsl:apply-templates select="//*[@id]"/>
    </duplicates>
  </xsl:template>


  <xsl:template match="*[@id]">
    <xsl:if test="count(key('ids', @id)) &gt; 1">
      <duplicate 
        id="{@id}" 
        dup-count="{count(key('ids', @id))}" 
        node-xpath="{string-join((for $node in ancestor::* return concat($node/name(),'[', count($node/preceding-sibling::*[name() = $node/name()])+1, ']'),concat(name(),'[', count(preceding-sibling::*[name() = current()/name()]) + 1, ']')
   ),'/')}">
     
      </duplicate>
    </xsl:if>
  </xsl:template>

</xsl:stylesheet>

Wednesday, 4 June 2014

Xsparql - a first attempt!

Xsparql is an easy method of using sparql queries with an xquery syntax. I thought I would give this a go with a data scrape from the schema.org markup on my own walks blog. The specific page I chose was The South West Coast Path - Marazion to Porthleven walk. I wasnt going to do anything complicated, just attempt to pull out the co-ordinates and names of the places featured in the walk posting.

First I needed an implementation of the xsparql specification, this I found at http://xsparql.deri.org/.

Then I needed an rdf source, this was created by scraping my blog posting mentioned above using Apache Any23 - Anything To Triples - Live Service Demo and requesting xml/rdf output which was saved as any23.org.rdf.

The xsparql code was then put together - it seemed to be white space sensitive in some instances and took a few attempts to get working. The code below was used and saved to a file named query.xs:

declare namespace place = "http://schema.org/Place/";
declare namespace geo = "http://schema.org/GeoCoordinates/";

<places>
{ for $Place $Name from <any23.org.rdf>
  where { $Place place:name $Name }
  order by $Name
  return <place obj="{$Place}" name="{$Name}" >
         { for $Name $Geo $lat $long from <any23.org.rdf>
           where { $Place place:geo  $Geo.
   $Place place:name $Name.
   $Geo geo:latitude $lat.
   $Geo geo:longitude $long}
           return <geo> 
   <lat>{ $lat }</lat>
   <long>{$long}</long>
  </geo>
         }
</place>
}
</places>

The implementation was then invoked with the following command line:

java -jar cli-0.5-jar-with-dependencies.jar query.xs -f result.xml

Which generateed the result document:


<places>
   <place name="Marazion" obj="b0">
      <geo>
         <lat>50.118267</lat>
         <long>-5.4776716</long>
      </geo>
   </place>
   <place name="Porthleven" obj="b1">
      <geo>
         <lat>50.118267</lat>
         <long>-5.4776716</long>
      </geo>
   </place>
   <place name="Prussia Cove Smugglers" obj="b2">
      <geo>
         <lat>50.101272</lat>
         <long>-5.4157501</long>
      </geo>
   </place>
   <place name="St Michael's Mount" obj="b3">
      <geo>
         <lat>50.116836</lat>
         <long>-5.4779291</long>
      </geo>
   </place>
</places>

Funky stuff!

Tuesday, 13 May 2014

Simple XSLT Construct Tester

When it comes to testing xquery constructs I find the easiest method is to use MArkLogics CQ interface. A simple answer to doing the same with XSLT is to use a single named template in an XSLT

Just create a basic XSLT named tester.xsl and add in one named template called 'testConstruct' and call it using the following command line

java -jar saxon.jar -it:testConstruct -xsl:tester.xsl -o:result.xml

The sample below tests creation of a string reconstruction from a test string

<xsl:stylesheet version="2.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
xmlns:xs="http://www.w3.org/2001/XMLSchema">

<xsl:template name="testConstruct">
 <xsl:variable name="provisions" select="'Sch. 08 para. 028(02) para. 058(03)A Sch. 15'" />
 <xsl:variable name="schtokens" select="tokenize($provisions, 'Sch. ')" />
 
 <xsl:for-each select="$schtokens">
  <xsl:variable name="paraTokens" select="tokenize(substring-after(., ' '), 'para.')" />
  <xsl:variable name="schNo" select="if (contains(., ' ')) then substring-before(., ' ')   else ." />
  <xsl:value-of select="if (not(matches(., '^\s*$')) and contains(., 'para')) then 
    (concat('Sch. ', $schNo, ' para.', substring-before($paraTokens[2],'('), 
     string-join(for $p in $paraTokens return replace(normalize-space($p), '^[0-9]+', ''),'')) )
    else if (not(matches(., '^\s*$'))) then 
    ( concat('Sch. ', $schNo) ) 
    else ()" />
  <xsl:if test="not(position() = last())">
   <xsl:value-of select="' '"/>
  </xsl:if>
 </xsl:for-each>
</xsl:template>
</xsl:stylesheet>

Wednesday, 16 April 2014

FOP External Graphics Caching

A recent issue encountered with FOP 1.0 involved the FOP processor grinding to a halt when a document contained a batch of external images hosted on a separate domain. It appeared to be fine with just two or three images but froze if greater numbers were included. The issue appeared to be centred around a prefetch utility in the image loader. This was attempting to fetch all the images in the document prior to processing and it appeared to be overloading the Apache Server with the request.

This was further substantiated when removing the images and then processing with just a couple, then reprocessing with the next few inserted. This worked because the initial ones had already been cached. However, such a routine is certainly not an efficient work-around.

The resolution was to disable the FOP image caching using a system property, -Dorg.apache.xmlgraphics.image.loader.impl.AbstractImageSessionContext.no-source-reuse=true which was added to the JAVAOPTS environment variable. This is easily done by modifying the FOP batch script that is part of the FOP download, adding the property to the appropriate line:

set JAVAOPTS=-Denv.windir=%WINDIR% -Xmx2048m -Dorg.apache.xmlgraphics.image.loader.impl.AbstractImageSessionContext.no-source-reuse=true

The result of this fix is that the processor only fetches the images when needed as seperate requests rather than as a bulk load in one request. The result worked perfectly. No more hanging, and even documents with 100's of images were being processed within a minute.

More information can be found at http://xmlgraphics.apache.org/commons/image-loader.html

Wednesday, 9 April 2014

Determine openssl version in apache

With the Heartbleed bug vulnerability around there is a need to determine the version on openssl used on apache systems

The status of the versions of Openssl affected are:

  • OpenSSL 1.0.1 through 1.0.1f (inclusive) are vulnerable
  • OpenSSL 1.0.1g is NOT vulnerable
  • OpenSSL 1.0.0 branch is NOT vulnerable
  • OpenSSL 0.9.8 branch is NOT vulnerable

How to determine the version of openssl that is being run in the apache installation on windows.

Open the command line and navigate to the apache/bin directory and use the following line

openssl version -a

To check openssl vulnerabilities in apache based sites use the online tool at:

http://filippo.io/Heartbleed