comment on

Same code with expanded regex with comments

#!/usr/bin/perl

# https://perlmonks.org/?node_id=1224600

use strict;
use warnings;
use Data::Dump 'dd';

my $data = <<'END';
interface XYZ
  given param1 -> child of "interface XYZ"
  given param2 -> child of "interface XYZ"
    given param2.1 -> child of "given param2"
      given param2.1.1 -> child of "given param2.1"
      given param2.1.2 -> child of "given param2.1"
    given param2.2 -> child of "given param2"
  given param3 -> child of "interface XYZ"
  given param4 -> child of "interface XYZ"
interface SECOND
  given param5 -> child of "interface SECOND"
END

my $struct = buildstruct($data);
dd $struct;

sub buildstruct
  {
  my $block = shift;
  my @answers;
  while( $block =~ /^ # make sure to start at beginning of a line (wit
+h m)
      (\ *)           # match leading spaces of header line
      (.*)            # match rest of line, save as head
      \n              # and match the newline
      (
      (?:             # match all following lines with
        \1            # same whitespace as head
        \ +           # plus at least one more space ( i.e. indented )
        .*\n          # contents to be looked at later
      )*              # as many as possible
      )               # save as rest
      /gmx )          # global, multiline, and extra whitespace
    {
    my ($head, $rest) = ($2, $3);
    $head =~ s/ ->.*//;
    push @answers, $rest ? { $head => buildstruct($rest) } : $head;
    }
  \@answers;
  }
[download]

I hope this helps :)

The regex matches each line, gets the indentation space string, then also matches all following lines that are indented that much plus at least one more space.

In reply to Re^3: Parsing by indentation by tybalt89
in thread Parsing by indentation by llarochelle

Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!

Titles consisting of a single word are discouraged, and in most cases are disallowed outright.

Read Where should I post X? if you're not absolutely sure you're posting in the right place.

Please read these before you post! —

Posts may use any of the Perl Monks Approved HTML tags:

a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, details, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, summary, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr

You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)

	For:		Use:
	&		`&`
	<		`<`
	>		`>`
	[		`[`
	]		`]`

Link using PerlMonks shortcuts! What shortcuts can I use for linking?

See Writeup Formatting Tips and other pages linked from there for more info.