Para leerlo en español, clic aquí.
Understanding Scientific Studies
Throughout our book Endometriosis: Putting the Pieces Together for a Better Life, we highlight details from various studies, especially in the chapters on treatment. We did this because we not only wanted to provide you with information, but we also wanted you to be able to read and understand some of the studies for yourself. We chose these studies for their relevance, but keep in mind that all of them have strengths, limitations, and biases that can affect the results and lead to inaccurate conclusions.
The most important thing to know about endometriosis studies is that many of them are of low quality due to their limitations and biases. This makes it difficult to draw solid conclusions from them. When reading a study mentioned in this book, you should never assume that its conclusions are an absolute truth, that they are certain, or that they apply to your situation.
Strengths, limitations, and biases
Strengths are the study’s positive aspects: what makes its results more reliable, such as good design, methodological rigor, precise measurements, an adequate sample size, and the use of valid and reliable instruments. They indicate in what ways the study is well-conducted and provides useful evidence.
Limitations are a study’s weaknesses: aspects that restrict the scope, validity, or generalizability of the findings, but do not automatically invalidate the study. They indicate the extent to which the conclusions can be applied and what should be interpreted with caution.
Biases are systematic errors that skew results in one direction, deviating them from the “truth” and distorting the conclusions. They are not mere random errors: they distort reality and present a false picture of the effects being studied. They can occur at all stages of the study (participant selection, methodology, measurement, analysis, publication), which reduces its validity and credibility.
Types of studies
Remember, not all studies are equally valuable. Some provide solid evidence, while others offer only initial clues. When a study is well-conducted (with few errors and good control of potentially influencing factors) it has a higher level of evidence. Therefore, levels of evidence help us assess how reliable a study’s results are.
From strongest to weakest
Systematic review and meta-analysis
Studies on the same topic can yield different and contradictory results. Systematic reviews are studies that gather, review, and summarize the available scientific evidence on a specific question or topic, providing us with a structured, qualitative overview. On the other hand, a meta-analysis uses statistical methods to analyze data from different studies to draw more robust conclusions and quantitative results. These types of studies are robust because they pool a large amount of data and reduce the error that can occur when analyzing a single study.
Studies on endometriosis tend to be highly heterogeneous in terms of their design and data collection, which can limit the conclusions that can be drawn from systematic reviews and meta-analyses.
Randomized controlled trial (RCT)
This study design is considered the gold standard in research. In RCTs, two treatments are compared to see how well they do. Participants are randomly assigned to either the treatment group or the control group. It’s considered a stronger level of evidence than observational studies because randomization minimizes bias.
If the RCT is blinded, the evidence is even stronger. Blinded means that participants don’t know whether they are in the treatment group or the control group. In a double-blind study, the healthcare professionals administering the treatment also don’t know which group the patients are in. In a triple-blind study, those analyzing the data are also blinded. This blinding technique is difficult to apply in some types of studies, such as surgical procedures.
Observational studies
This is a study design in which different groups are followed over time without being randomly assigned to a treatment; researchers simply observe what happens to them. This category includes cohort studies and case-control studies.
- A prospective cohort study follows groups over time.
- A retrospective cohort study analyzes available historical data for the groups.
The major drawback in both cases is that observational studies suggest associations but rarely can demonstrate causality with a high degree of certainty. It’s very important to interpret the results with caution.
Case reports and case series
These are the simplest and most descriptive designs in scientific evidence. They are detailed accounts of one (report) or several patients (case series) that help identify new or rare problems. However, by their very nature, they are not sufficient to confirm that a treatment works or that something causes a disease. That’s why we need to look for larger studies with more scientific evidence to verify whether something works or not. Their value lies in enabling researchers to recognize rare signs and symptoms and generate questions for larger studies (mentioned above).
General information and expert opinion
These include all background information (books, manuals) and recommendations or judgments made by individuals with extensive experience and knowledge in a clinical area. Although they are useful for understanding a problem and for guiding us when there are no good studies available, they tend to be more general and based on experience. In many cases, the information in books or provided by our healthcare professional is not up to date or based on a rigorous review of the latest information. Therefore, we must approach it with caution and rely on more solid evidence.
Key findings in studies
Control group
A control group is a group in a study that doesn’t receive the treatment. The control group may take a placebo or may undergo expectant management. Expectant management doesn’t involve taking a placebo; participants are simply monitored to see if their condition improves, worsens, or remains unchanged, using a “wait-and-see” approach.
For example, an RCT might compare dienogest with a placebo (control group). A retrospective cohort study might compare data from the group that took dienogest with those from the group that did expectant management (control group).
Studies don’t always have a control group, as they sometimes compare treatments with one another, such as dienogest versus oral contraceptives.
“Significant” results
When results reach statistical significance in a study, this means it’s more likely that the results are due to the relationship between the two factors rather than chance. For example, when a study finds that surgical excision of intestinal endometriosis significantly improves pain, this means that the statistical analysis found that the improvements in pain were likely due to the surgical excision and not just chance.
Average data
Studies often report average data, such as the average age of participants or the average follow-up time, but remember that an average is calculated by adding up all the data and dividing it. The range of data is usually much wider than the average. For example, the average follow-up time might be 3 years, but researchers followed some participants for 5 months and others for 6 years.
Limitations
All studies have limitations. Here are some examples:
Missing or incomplete data
This occurs when not all data from all participants is available. This can be especially problematic in retrospective studies. In most studies, some participants drop out partway through the process, so the rate of participants lost to follow-up (and who consequently lack complete data) can affect the results.
Recall bias
This occurs when patients have to recall and report on something from the past and, therefore, may not remember it accurately. For example, questionnaires sent to patients asking them to recall data (such as how often they experienced fatigue over a 6-month period) carry a higher risk of recall bias.
Follow-up period
This refers to the duration of the study. For example, some studies on surgery only analyze outcomes regarding pain levels between 6 and 12 months after the procedure. Given that endometriosis is a chronic condition, these are short-term data that may differ if we analyze the results three years after surgery. However, longer studies are more expensive, carry a higher risk of participant dropout, and require more time. For this reason, many long-term studies on endometriosis are retrospective cohorts, meaning they analyze data from two groups rather than following them prospectively.
Sample size
This refers to the number of patients in a study group. Generally, the larger the sample size, the greater the statistical power to analyze the data and determine whether the result is significant and not merely coincidental. The results of a well-conducted study with 500 participants are more robust than those of the same study with only 50 participants. However, due to the cost of conducting studies with many participants, large studies are often retrospective studies that analyze existing data records.
To obtain larger samples, it’s common to conduct the same study at multiple centers and pool the data. This usually works well in drug studies, but in surgical studies it can negatively affect the results, since each surgeon has different levels of skill and experience that influence the outcome.
Study design
If a study has a design flaw, it can lead to inaccurate conclusions. For example, a study comparing oral contraceptive pills with excision/ablation surgery would have a design flaw, since these two types of surgical procedures are very different and cannot be grouped together.
A brief summary of what we’ve learned
Studies provide different levels of evidence depending on their type, design, sample size, follow-up periods, biases, and limitations. Many studies on endometriosis are of low quality, making it difficult to draw solid conclusions from them.