Saturday, 26 September 2015

Some problems with the Wilkins Rate of Reading Test


The figure above shows a visual fields test that displays the sensitivity of the visual field at locations around the fixation point. The black areas represent reduced sensitivity. You might think that there has been some improvement in vision as a result of some unspecified intervention. However, if you tried to publish a figure such as this,  a good peer reviewer would spot the problem straight away.
Before you can say whether the result of a psychophysical test has improved you need to know if there is been some change in the way the observer is performing the test. Are they performing the test in a more or less conservative manner for example? For this reason, machines for testing visual fields build in catch trials to determine if the subject is becoming a more or less conservative observer. For example, they make noise as if a stimulus has been presented when it has not, or they may measure the stimulus response time - anything less than 200ms is almost certainly a false positive. Going back the case above, the tests showed that the observer had become a less conservative observer or more trigger happy if you like and this was the most likely explanation for the improvement int he visual field see below.

So, what has this got to do with the Wilkins Rate of Reading Test (WRRT)? It is important to
remember that the WRRT is not a standardised reading test. It not used by anyone except a
aficionados of visual stress and its treatment. The WRRT does not consist of naturalistic text but closely spaced, commonly used words, ordered randomly without syntax, punctuation or paragraphs. Its proponents argue that it is well suited to detecting visual problems. That may be true, but it also has some key flaws. The principle one being that it can not measure 'response criterion' That is, has the observer become more or less conservative in his responses? It can not be said to isolate out visual factors if you can not measure other aspects of reading. For example, you might find that subjects were able to read faster but at the expense of comprehension, parsing or miscues.
I know this from my own experience of learning Spanish through skype. When asked to read text I can easily vary my rate of reading making it sound fast and 'spanishy'. However, my teacher notices the parsing goes a bit funny and she throws in a comprehension question. The rate of reading then comes crashing down. Although she wouldn't call it that, she has recognized that my response criterion has changed and I have become a less conservative reader - taking risks with pronunciation, sentence structure, and meaning.
So, returning to the WRRT unless you can show that other aspects of reading are conserved you can not claim to have isolated out visual aspects of reading. Also, you can not claim that any improvements in the WRRT have any relevance to real world reading.
Still, you have to admire somebody who has found a way to make money out of printing random words.

Sunday, 6 September 2015

Intuitive Colorimeter: Technique for the 21st century?


The intuitive colorimeter is a device that is said by its proponents to provide the best means of prescribing colour to alleviate visual stress. Kriss and Evans for example state that ‘people almost invariably report more benefit from precision tinted lens than from coloured overlays because precision tinted lenses are easier to use (eg with white boards and when writing) and because colour can be prescribed with more precision’(1)
In a lecture hand-out, Bruce Evans has argued that the Intuitive Colorimeter Mark III is the 21st-century method for assessing visual stress. So I thought it would be worth assessing the evidence for this device.
Random letters (why?) illuminated in the Intuitive Colorimeter
The intuitive colorimeter is, in essence, an illuminated box with an aperture for viewing text. The hue and saturation of the light in the box can be changed. The light source is a halophosphate fluorescent tube. Rotating the dial alters the filter condition to change the illumination of the text which consists of random letters (why?) The colorimeter setting can then be assessed against a set of trial lenses which are then be prescribed in the form of glasses - so called precision tinted lenses.
So what is the evidence for lenses prescribed with the intuitive colorimeter?
I have been able to identify a hand-full of trials(2)(3)(4) (5)(6) which address this issue. Taken together or individually they do not amount to a compelling case for the use of the intuitive colorimeter.

The trial by Wilkins and Colleagues published in 1994 in the Ophthalmic and Physiological Optics in 1994 is often cited in support of precision tinted lenses even though it is a negative or at best inconclusive study. 68 poor readers with visual stress, diagnosed according to the criterion of voluntary sustained use of overlays, were enrolled. All subjects were tested with the intuitive colorimeter and their preferred tint identified. They were also prescribed a closely related tint which did not ameliorate their perceptual distortions.
The study had a crossover design and subjects were randomised to use experimental tint or placebo tint first.
Unfortunately, the dropout rate from the study was so high that it is not possible to draw meaningful conclusions. For the 45/68 for whom reading data was available, there was no difference in reading speed accuracy or comprehension comparing preferred lens with placebo lens. Interpretation of the symptom diaries is even more problematic because data was only available for 36/68 subjects. Using experimental lens 71% of days were symptom-free and with the placebo lens 66%. An unimpressive set of results when you remember that nearly 50% of the data is missing 
Perhaps the most striking result is that 31 preferred the first lens and 17 the second; a difference that is highly significant and suggestive of novelty or Hawthorne-effects.

The trial by Mitchell and colleagues published in 2008 was a parallel groups study with a control group of 17 who received no lens, a placebo lens group of 15 and an experimental lens group of 17 who received their chosen tint as determined with the Intuitive Colorimeter.
The subjects were individuals with reading difficulties who complained of visual distortions so the diagnosis of visual stress was purely symptom based.
The outcome measures were symptom scores from the Irlen Differential Perceptual Scale and the Neale analysis of reading Test. There was no difference between the experimental lens or placebo lens groups for any of the measures although there were differences with the no treatment control group which is suggestive of a strong placebo effect.

Singleton and Trotter 2005 Like many studies in this area, this study starts with a small group which is divided into even smaller groups. As a result, the statistical power of the study is limited.
There was a total of ten students with dyslexia of whom 5 had visual stress and there were ten normal readers of whom 5 had visual stress diagnosed according to the Visual Processing Problems Inventory VPPI. There was no placebo control group. The study was a crossover design that compared lenses prescribed with the intuitive colorimeter against no lens. The outcome measure was the Wilkins Rate of Reading Test.
The only group to show some improvement were the five patients with dyslexia and visual stress  However there are a number of serious problems with this study which include. Small sample size and no placebo control group.
The external validity of this study is poor because subjects were recruited via the disabilities service of the University and included subjects who may have been aware of the potential benefits of colour. Finally, because the WRRT is a non-standardised reading test and it is not clear if the results can be generalised to naturalistic text.

Machlachlan  Yale and Wilkins 1993 wasn’t a controlled trial at all. Fifty-five subjects were recruited. Twenty-three of the cases were volunteers responding to an article in 'Living magazine' and the remainder were referrals from educationalists, psychiatrists and neurologists. As a result, external validity is likely to be low.
The criteria for diagnosing visual stress was awareness of visual distortions, sustained use of overlays and having chosen a colour with the Intuitive Colorimeter.
The outcome measure was what percentage continued wearing glasses over 10 months. 82% reported that they were still using the glasses 10 months later. 
Unfortunately, this based on self-report and is unreliable. The external validity of this study is questionable because of the way the subjects were recruited and finally.. so what?

Simmers et al 2001. This study comes from a reputable psychophysical laboratory at Imperial College. It has been reviewed in a previous post. Instead of using text it looks at the kind of geometric patterns that are said to be most aversive in visual stress. The authors found no difference between subjects with visual stress and controls. In those cases with visual stress no difference in aversive symptoms when using precision tinted lenses.

So -Intuitive ColorimeterTM method for the 21st century? If you believe that you will believe anything.


1.         Kriss I, Evans BJW. The relationship between dyslexia and Meares-Irlen Syndrome. J Res Read. 2005 Aug;28(3):350–64.
2.         Wilkins AJ, Evans BJ, Brown JA, Busby AE, Wingfield AE, Jeanes RJ, et al. Double-masked placebo-controlled trial of precision spectral filters in children who use coloured overlays. Ophthalmic Physiol Opt J Br Coll Ophthalmic Opt Optom. 1994 Oct;14(4):365–70.
3.         Mitchell C, Mansfield D, Rautenbach S. Coloured filters and reading accuracy, comprehension and rate: a placebo-controlled study. Percept Mot Skills. 2008 Apr;106(2):517–32.
4.         Singleton C, Trotter S. Visual stress in adults with and without dyslexia. J Res Read. 2005 Aug;28(3):365–78.
5.         MacLachlan A., Yale S., Wilkins A. Open trial of subjective precision tinting: A follow-up of 55 patients. Ophthalmic Physiol Opt. 1993;13(2):175–8.
6.         Simmers AJ, Bex PJ, Smith FK, Wilkins AJ. Spatiotemporal visual function in tinted lens wearers. Invest Ophthalmol Vis Sci. 2001 Mar;42(3):879–84.

Thursday, 3 September 2015

A big problem for the visual stress hypothesis

Spatiotemporal function in Tinted Lens Wearers
Anital Simmers, Peter Bex, Fiona Smith, and Arnold Wilkins
Invest  Ophthalmol Vis Sci 2001;42:879-884

For once, a rather good study which comes from a reputable group at Imperial College London. You can download it here.
Proponents of visual stress as a factor in reading difficulties argue that the sort of discomfort and movement illusions that most of us experience in response to recurring geometric patterns, varies in the population from person to person. They go on to argue that in some susceptible individuals it can be brought on by print, resulting in distortions and movements of letters and words - so called visual stress.
The conditions that induce these symptoms, for geometric patterns, are said to be surprisingly specific and peak with high contrast square gratings between 2-8 cycles per degree. Proponents argue that text has spatial frequencies within this range see the figure below.
The text has been filtered to remove high spatial frequencies and the contrast exaggerated to make the stripes more apparent
Text obviously isn't a striped grating. However, like a piece of music which contains multiple frequencies of sound waves at the same time (which you don't notice as the music washes over you), text contains multiple spatial frequencies and it is those between 2-8 cycles per degree that are the said to be a problem in 'visual stress'.
If this is the case we would expect to see this effect in simplified settings using actual gratings rather than text. Such a study has been done on subjects with visual stress, in a reputable psychophysical laboratory. Even more interesting is that Arnold Wilkins was one of the authors although he does not emphasise the results of this study a great deal.
The subjects were twenty individuals with 'visual stress' who had successfully worn tinted lenses (prescribed with the intuitive colorimeter) for at least six months. Twenty control subjects were recruited from the staff and siblings of staff and students of the University of Exeter.
Subjects were tested with a range of psychophysical tests which are designed to assess visual processing in the range of frequencies that are said to be aversive in visual stress.  The participants with visual stress, they were tested with and without their lenses.
The tests included spatiotemporal contrast sensitivity, contrast increment thresholds, random dot motion coherence, and motion perception
Results
To put it simply, there was no difference between control subjects and visual stress subjects and just as important, in the group with visual stress there was no difference with and without their lenses.
Contrast sensitivity function for subjects with visual stress with without lenses
Despite this, some people (mostly with a vested interest) still claim that lenses prescribed with the intuitive colorimeter are the 'gold standard.
Will this study dent their confidence? I doubt it.

Saturday, 22 August 2015

A 'trial' that doesnt quite 'stack-up'

Proponents of colour to treat visual stress in poor readers frequently refer to this study. To me, it looks very much like a post-hoc data trawl, or exploratory analysis, which is at best hypothesis generating.
The study is poorly written and it is hard and sometimes impossible to extract the raw data.

Tyrrell R, Holland K, Dennis D, Wilkins A.Coloured overlays, visual discomfort, visual search and classroom reading. JRes Read. 1995 Feb;18(1):10–23.

A total of 60 children were included (see below) however the main experimental group consisted of 46 children who were categorised into, above average readers (10), average readers (18) and below average readers (12). An additional group of older children age age 14-16 who were well below average readers were also identified (6)


 Two control groups were also identified. They were matched with the 6 well below average readers in terms of reading age (RA control group) and Chronological age (CA control group). Unfortunately when it comes to the results, the CA and RA control groups are not compared with the well below readers but with the whole group rendering their RA and CA control status invalid.

 Procedure
  Subjects were tested on three occasions and that testing must have been pretty long and tiring.
It involved testing from the scotopic sensitivity screening manual, choosing overlays, reading for 15 minutes with or without an overlay and a visual search task. When you consider that the key outcome was slowing of reading in the final five minutes of a fifteen minute reading task, it would be useful to know how long this whole process took.



Results.
So how about the results.  Looking at table two below it can be seen using the criterion of immediate benefit from an overlay, 100% of well
below average readers 75% of below average readers, 56% of average readers and 40% of above average readers had visual stress. That looks rather high to me. It is important to remember that screeners were not blinded to the status of the readers. So room for bias here.
There is also some surprising data resulting from Irlen's 1983 tests of perceptual difficulty. 75% of above average readers had moderate perceptual difficulty! And almost nobody had low perceptual difficulty.

The key finding of this study and one that does not really stand up to scrutiny is shown in the table below. Looking at the column on the left that is really a crossover study comparing performance with chosen overlay and no chosen overlay. It is at high risk of bias because
there is no placebo used and external validity is low because of the circumstances of testing; a fatiguing session at which a number of different variables were tested. The results show no overall improvement with overlays, but without overlays there was some slowing in the final five minutes. Please note however that this is expressed in syllables per minute not words. Even if you believe that this study is methodologically sound (which I do not) the results in unlikely to be educationally significant. The difference in words was very small indeed. Also note that the reading age controls are clearly not comparable with the group who chose overlays.

Summary.

This study is at best an exploratory study and because of what appears to be a flexible post hoc data analysis this study is at most hypothesis generating.
There are serious problems with 'internal validity'.
1) No placebo control group  -compares coloured overlay with no overlays
2) Post hoc data analysis - no pre-trial protocol available
3) Non standardised reading test expressed in syllables per minute which exaggerates any effect
4) No overall improvement with overlays. Possibly got less worse in the final 5 minutes.
5) No proper comparisons made with the CA and RA control groups
6) Small groups and low statistical power

Problems with external validity
Although subjects were recruited from a school setting which is good, the main problem was that reading was assessed in the same session as the Irlen perceptual tests, the overlay selection procedure and visual search task. It is perhaps not surprising that some subjects fatigued during the final five of 15 minutes reading aloud.

Overall another epic failure of peer reviewing!






 

What is medicine's 5 sigma?

Another piece exploring the same theme as the previous post - that is the poor quality and high risk of of false positive results in much of the scientific literature. This time my source is an editorial from the Lancet a leading general medical journal. You can read it all here

Richard Horton the editor of the Lancet attended a meeting at the Welcome Trust in London. The theme of the meeting was the current state of the biomedical research literature and the endemicity of bad research. As one speaker put it - poor methods get results. Another stated that too any scientists sculpt their data to fit their preferred theory of the world or retrofit their hypotheses to fit the data (see the previous post about arrows and targets)
The current estimate is that about half of studies contain false positive results. So when proponents of the use of colour to treat visual stress ( and any number of other conditions) say that research in the peer reviewed literature has shown that ............ ( you can fill in the blank) that is not enough.These studies have to be looked at in detail.

A number of factors were identified that put studies at high risk of producing false positive associations. These were

1) Small sample sizes
2) Tiny effects
3) Invalid exploratory analysis
4) Flagrant conflicts of interest
5) Pursuing fashionable trends of dubious importance

Much of the research on visual stress in dyslexia ticks all of these boxes. While this does not remove the need to read each study in detail and with  crtitical eye the statement that research in the peer reviewed literature has shown that........... is not nearly as impressive as it sounds.
Onwards with more reviews

Sunday, 16 August 2015

More negative studies and why it is a good thing

An important study has appeared in the journal PLoS One  that shows that fewer trials in the field of cardiovascular medicine are reporting positive outcomes. Why is this good new? And what does it have to do with the treatment of visual stress?

First; some background.  As many as 50% of published studies contain false positive results and the situation is particularly bad in the field of neuroscience and psychology. While some false positive results are inevitable, the numbers at present suggest some systematic biases in the literature.
For example, in the field of fMRI studies more positive findings are reported that the study designs can support. Trials have certain power or ability to detect  significant differences which depends on the sample size and variability of the population being studied. John Ioannidis has shown evidence of  bias (in the statistical sense of the word) operating. Small studies with a lower power to detect signal changes are finding as many positive findings as larger more powerful studies. Naturally, you would expect smaller studies, with less statistical power, to find fewer positive effects.You can download the study here.
So how has this state of affairs come about? Not through fraud as you might think -although that can happen.  Human nature and the subjective biases that affect us all are probably the culprit.
The first problem is that studies reporting positive outcomes are more likely to be published; so called publication bias. One of the things you have to do when reviewing the literature in a systematic fashion is to look out for unpublished material which may not have found a home because negative findings were being reported.
Another important factor is a flexible approach to data analysis. If you do an exploratory study in which you measure multiple variables, and analyse your data in multiple different ways you are quite likely to find some positive results which pass the arbitrary criterion for statistical significance of 0.05. For example you could stratify your groups multiple different ways or you could have multiple endpoints and not declare which was the primary endpoint.
This has been compared to the man driving past a barn who sees  lots of targets with arrows in the centre and assumes the farmer must be a pretty good shot. Then, as he drives a little further he sees the another wall of the barn  where the farmer is painting targets around arrows he has already fired into the wall! It's a bit like that with trials if you allow a flexible post-hoc approach to data analysis.
This may be acceptable for early exploratory studies within a field but such studies are at best hypothesis generating and the results need to be confirmed by properly conducted RCTs.
To get round this problem many funding bodies  insist that researchers pre-register the trial, stating what  what the outcome measures will be,  what subjects are going to be studied and what statistical tests will be used. This, in crude terms, is equivalent to painting the target on the wall before the arrow is fired. With this has come a reduction in positive findings and that is a good thing. False positive studies waste resources and can endanger human life.
So what has this got to do with the treatment of visual stress? Well, many of the studies of treating 'visual stress' in dyslexia show the hallmarks of a flexible approach to data interpretation, Dividing the subjects into multiple small groups and studying lots of different outcome measures. Then pouncing on those that appear positive an ignoring the rest.
Before anyone starts to take these ideas seriously we need a randomised controlled trial with pre-registered protocols and outcome measures.


Wednesday, 29 July 2015

Postscript to -The trials they don't want you to know about (1)

As a follow up to their paper published in 2011 Ritchie and colleagues published a study looking at the subjects of their first trial, one year on.

Ritchie SJ, Sala S Della, McIntosh RD. Irlencolored filters in the classroom: A 1‐year follow‐up. Mind Brain Educ. 2012Jun;6(2):74–80.

Though not as rigorous as their previous study, outlined in my last post, it does contain some fascinating data that does nothing to support the use of of colour to treat the overlap group with poor reading and visual stress. I suspect it was a bit of an afterthought put together following criticisms of their first study - that benefits would not be seen without longer term follow up.
In the first study an Irlen screener diagnosed 47 out of 61 poor readers with 'visual stress'.
Those 47 children were then tested with prescribed overlay and placebo overlay and no difference in the rate of reading using the WRRT or reading naturalistic text measured with the Gray Oral Reading Test (GORT).
After 12 months 22 out of 47 of the original sample were still using colour in some form and 18 of these were still available to study. So this was a highly selected group who passed two criteria for visual stress; the Irlen screening tests and voluntary sustained us of overlays. According to their teachers this group had used overlays or lenses for a substantial part of the time. For this reason, if coloured overlays are effective, you would expect to see real world improvements in reading in this group.
Unfortunately the controls were a group of poor readers without visual stress from the previous study. The authors ackowledge that the ideal control group would have been individuals with visual stress who had been treated with a placebo overlay.
The visual stress group was tested with the WRRT using prescribed filter, placebo filter and clear filter. In short this part of the study was a crossover or within subject design.
The other test used was the Gray Oral Reading test which is a standardised reading test that measures accuracy and rate of reading and crucially, questions to test comprehension are included. The outcome is standardised for age to produce the Oral Reading Quotient (ORQ). This means that  if a subject progresses normally for age the ORQ should remain the same. If there is a catch up in reading, as predicted by proponents of the treatment of visual stress, the ORQ should improve and then perhaps stabilise.
Results
 The WRRT. No surprises here really. There was no significant difference in reading rate between presecribed overlay, placebo overlay and clear overlay.












GORT Oral Reading Quotient ORQ
Figure two on the right shows that for both groups, the Irlen wearers and the controls, the ORQ declined over the study period. This occured in spite of imrovements in WRRT for both groups bringing in to doubt the validity of WRRT as a reading test.



Conclusions
The authors are good scientists who are careful not to over-interpret their study and they also acknowledge its weaknesses. In particular the small sample size and the lack of a visual stress control group. Nonetheless it is one of the few studies to take a longer term look at outcomes in a group where a difference might be expected to be found. That is, those with visual stress diagnosed by an Irlen-screener and who pass the criterion of voluntary sustained use over 12 months. In spite of perceived subjective benefit no improvement in real world reading was seen.